Non-programmers Were Sneaking Into Codex, and ChatGPT Work Followed
OpenAI handed knowledge workers the Codex harness essentially unchanged, altering only the UX trade-offs and the sandbox defaults. The judgment underneath: AI will erase the boundaries between job functions, so OpenAI refuses to draw product lines around "who you are".
The video won't play here. Listen to the audio instead:
The argument · tap a timestamp to hear it
Non-developers sneaking into Codex is where this product actually starts
Akshay says the most immediate trigger for ChatGPT Work was an inflection point inside OpenAI: non-developers started adopting Codex — strategic finance, marketing, all of them using it. But the signal that stuck with him wasn't usage volume, it was the emotion. These people were "very proud" of the fact that they were using Codex. Swyx named it on the spot: "I'm not supposed to be using this, but I am." Akshay agreed that was exactly it, and added that they felt they had acquired a superpower. The conclusion he draws: the moment when agent capability spills beyond developers arrived far earlier than OpenAI itself expected, and ChatGPT's existing distribution is a ready-made channel. All that remained was figuring out how to put that power in front of those people.
— Akshay NathanBoth products share one harness; the difference is in the defaults
Swyx asked directly whether Codex and ChatGPT Work run on the same harness. Akshay's answer: the same one, shared — and the improvements built for knowledge work (plugins, computer use, artifacts) are available on both sides. There are exactly two differences. First, UX trade-offs: Codex mode pushes Git state and diffs to the foreground by default, the dynamic island assumes you're inside a repo, and the same "build me a retirement calculator spreadsheet" task shows no file edits at all in Work. Second, the sandbox has different security defaults. His governing principle: a user should not be forced to pick which experience they're in before they start. And the reason for merging wasn't engineering reuse — it's that every few months everyone's work changes, so drawing hard boundaries by identity is bound to be wrong eventually.
— Akshay Nathan32 options is too many; the defaults are the product
Asked how to min-max between models and reasoning tiers, Akshay leads with a number: 32 options — and volunteers that "arguably there are too many options right now, and we're simplifying." The ordering he gives is firm: use the default first, and only reconfigure once you see a specific unmet need on cost efficiency or on intelligence quality. Ultra and the multi-agent settings suit two kinds of task — extremely complex open-ended exploration, or highly parallelizable work; goal suits tasks that can make sustained, verifiable progress; most tasks are neither. The slider is designed to project a multi-dimensional space onto one dimension by force: speed and efficiency at one end, quality and thoroughness at the other. After launch they also changed Ultra so you have to turn it on manually in advanced settings, because it burns through usage limits faster.
— Akshay NathanSites are replacing the deck-plus-spreadsheet
Most of the discussion around Sites focuses on their prototyping side; the other side Akshay describes deserves more attention. The company's finance team produces a monthly report that historically was a slide deck plus a spreadsheet. Now it is simply a Site, and the Site itself has become the vehicle for team collaboration. His reasoning is bandwidth and ceilings: PowerPoint and Excel are nominally infinitely flexible, but you hit one of two walls — either the person doesn't know how to use some feature, or the product simply doesn't support it. A Site is HTML, so if you want a particular conclusion to be the hero image, it can be; things that would get buried inside a deck can be pushed to the top. Even the model slider in this very launch was finalized by design, engineering and product collaborating on a Site. Akshay's verdict on Markdown: it was never meant for humans to read.
— Akshay NathanAI doesn't write performance reviews; it goes and finds all the context
Swyx's concern is concrete: as a manager, using LLM-generated output to evaluate a person makes it look like you don't care. Akshay draws a clean line — he would never rely purely on AI to write a review and then submit it. The value of AI is entirely in gathering context: the code, the problems this person caught, review records, Slack — it can go through all of it. He says he tried this in the review cycle six months ago and it was no help at all; this cycle the model did better than he did on his own, surfacing contributions he simply hadn't seen. Swyx summed it up as "this is just agentic search", and Akshay agreed, adding that it's now more steerable and more targetable than before. As for whether it misses things: yes, he says, but so do I — it's relative.
— Akshay Nathan10 million users is not an achievement, it's a numerator
Akshay calls 10 million users a culmination, then immediately puts it back over the denominator: ChatGPT has hundreds of millions of users, nearly all of whom equate it with AI, and against that 10 million is only a start. The binding constraint is one OpenAI imposed on itself — ChatGPT Work is currently open only to paying users, and OpenAI will not switch ChatGPT users into it by default. It has to come through education, trials and rounds of feedback and iteration. He also explicitly rejects the idea that Codex will be absorbed, in stronger terms than Swyx's question: developers have always been the core market, Codex will keep going deeper on software development, and the merge in fact raises its utility, because you can move seamlessly between writing diffs, generating artifacts and searching across codebases.
— Akshay NathanThe bottleneck has shifted from execution to running out of new ideas
When everyone can build things, Akshay argues, the bottleneck becomes ideas and taste. He says the one automation he most wants and still cannot get is "give me new ideas" — LLMs are simply not good at it. His explanation isn't that models are insufficiently capable; it's the structure of where ideas come from. Ideas don't emerge in a vacuum. In product development they come from talking to users, from the friction and feedback you've observed, from the groundwork you've already laid. So he expects generalists who can close that loop to remain valuable indefinitely. On how roles evolve, his view is T-shaped: AI turns everyone into a generalist — he says he used to be unable to produce design at all and still has no visual taste, but he can iterate on it with AI — and the deep specialty is the vertical stroke.
— Akshay NathanDon't mistake adding dashboards for a team making progress
The old proxy metrics for productivity — commits, lines of code, story points — are breaking down in the AI era: token usage and PR counts are no longer strongly correlated with whether a team is hitting its goals. What Akshay offers managers instead is not a new dashboard but at-bats: not just the count, but the quality — whether the full loop runs efficiently, from generating an idea, to building it, to getting feedback, to reacting to that feedback, to validating or falsifying the hypothesis, and on to the next idea. And whether the team has the kind of culture that keeps running that loop with motivation and humility intact. He names the biggest trap explicitly: you've added a pile of LLMs and built all kinds of dashboards, but nothing has changed — that is exactly mistaking motion for progress.
— Akshay NathanIn their own words · checked verbatim
And, like, we should enable users to choose, but we shouldn’t box them in.
Akshay Nathan11:43
at these tools like PowerPoint and Excel are like infinitely flexible, but at some point you reach the boundary of like either as a human you may not know how to use some feature or something, or the product itself doesn’t support it. But with a site you can do anything.
Akshay Nathan23:52
Oh, I think the etiquette is that, like, I would never write something via, like, well, solely via AI and, like, present it as, like, a review for someone. What I was talking about is more, like, gathering context. That’s the place where it’s incredibly helpful.
Akshay Nathan38:27
I would say that the one automation that I would love to work and it doesn’t work is bring me new ideas, right? somehow LLMs are just not it. One interesting part about ideas is, like, they’re not, like, in a vacuum.
Akshay Nathan1:03:07
I think maybe the trap is like conflating motion and progress. I think motion is much easier now than ever before because of the tooling that we have. But progress requires you to be like very prescriptive and deliberate about like what you’re trying to achieve
Akshay Nathan1:08:06
Figures
| OpenAI headcount when Akshay joined | About 500 people (2023) | 1:33 |
| Model and reasoning tier combinations | 32 options | 16:24 |
| ChatGPT Work + Codex users | 10 million | 41:43 |
| Total ChatGPT users | Hundreds of millions | 42:46 |
| ChatGPT Work availability | Paying users only, and not enabled by default for ChatGPT users | 42:46 |
| Same task, time compared (retirement calculator) | 8 minutes on ChatGPT Work; Codex still hadn't finished at that point | 19:29 |
| Time for Vibhu to generate the same board game with goal | 18 minutes 53 seconds | 31:19 |
Glossary
- harness
- The layer of tools, loops and environment wrapped around a model, determining what an agent can call and how it executes.
- artifacts
- Structured outputs a model generates and can iterate on repeatedly — spreadsheets, slides, web pages.
- Sites
- HTML pages ChatGPT generates and hosts directly, now replacing decks and spreadsheets as deliverables.
- Chronicle
- An experimental feature that learns from how you use your computer and writes to memory. Off by default.
- at-bats
- A baseball metaphor: the number and quality of complete idea-build-feedback-validate loops a team runs.
- OpenClaw
- A standalone open-source personal agent project, used in this episode as a reference point for the personal-OS form factor.
How to listen
PMs and designers building agent products; anyone rolling out AI tooling to non-engineering teams; engineering managers who need to redefine what productivity means on their team.
26:02-31:19, where the two hosts take turns demoing the board-game Sites they built. Low information density.