Operator Is Not a Continuation of O3 — It's the First Hand of Reasoning Reaching Into the Physical World
OpenAI tried a Web Agent in 2016, failed, and laid off twenty or thirty people; what was missing was a foundation model. Ten years later, Operator lines up three things — a foundation model, human data, and reinforcement learning — and only then does reasoning move from abstract text toward visual interaction.
The video won't play here. Listen to the audio instead:
The argument · tap a timestamp to hear it
Operator is closed-loop control; O1 is open-loop
Wu Yi defines Operator's place on the AGI roadmap as ‘a general-purpose shift’: O1 and O3 are transformations that think on a focused problem, but there is no environment interaction in between, and no multimodality. Operator fills in both — it can see an observation, take one action, then see a new observation, doing many rounds of interaction in a row, forming a closed-loop control system. Traditional large models are open-loop: give an instruction and it outputs directly, with no feedback wanted; Operator, after outputting, actively makes a call and expects the next feedback, then thinks about the next step based on that feedback. This ‘horizontal-axis’ general capability is something you must have on the way to a general agent.
— Wu YiOperator's chain of thought is short; backtracking happens on actions
Wu Yi does not think Operator is a simple continuation of O1 and O3. Observing the chain of thought Operator released, he finds it relatively short, and backtracking mainly appears on actions — for example, it clicks a webpage and it doesn't open for a long while, so it clicks back. O1 and O3's chains of thought are very, very long, with lots of backtracking. So he leans toward thinking Operator is not as extreme a model as O3, more like ‘a multimodal closed-loop O1, an Agent O1 version’, or possibly a spot a bit more O1. It moves wider in breadth rather than going to the extreme in depth, so there is still room.
— Wu YiReplicating Operator takes three things, and China isn't that far off
Wu Yi breaks Operator's technical essentials into three: a good native multimodal foundation model, high-quality human data and tasks, and an efficient large-scale reinforcement learning system that supports agents. The third is the most troublesome, because in an agent environment the model spits out several hundred tokens and then has to click on the screen; you need a computer sitting there waiting for the AI to click, doing simulation, and how to interact efficiently at the scale of several thousand or ten thousand cards is a research problem. But he judges that Chinese teams catching up to Operator is not as hard as catching up to GPT or Sora: the route is clear, a lot of the infrastructure is already there, and it is a half there state. The problem is you don't know how much OpenAI has underwater.
— Wu YiOpenAI's first big project was a Web Agent, and it failed
Wu Yi goes back to 2016: the first big project OpenAI did after being founded was a Web Agent, a general visual agent clicking on webpages. At the time they only had reinforcement learning, no foundation model, and nobody even went to find data annotation; there wasn't even transmission. They used a big LSTM called Covnet to click webpages, clicked with reinforcement learning, and then failed, and laid off that team, laying off twenty or thirty people. Wu Yi tells this story to students in his deep learning class, saying the missing recipe at the time was the foundation model — reinforcement learning alone won't do, and a good foundation model alone won't quite do either; you need the two together. Ten years later they got it done.
— Wu YiLevels three to four is a huge change, not just two or three years
Wu Yi unpacks OpenAI's five-level classification: a chatbot is reactive, you say a line and I say a line; reasoning is thinking in your head for ten seconds before speaking, with reasoning and planning; an agent adds an external world, a three-party process, and the AI's scope gets bigger. But levels one through three are all instruction following or instruction execution: give an instruction, complete it, and you can verify right or wrong. Innovation is different — you can only give something directional, and right or wrong is meaningless; what you want is something good that goes beyond the original knowledge system. Wu Yi thinks three to four is a huge change, possibly not achievable in just two or three years. And which of four and five comes first isn't certain; organizations may form passively — in the future every piece of software will carry an agent, and interactions will exist between them.
— Wu YiMulti-agent interaction won't appear any time soon
Wu Yi's view is that for the time being, situations where multiple agents interact won't appear. Because if the goal is to complete a task, one sufficiently general large model can do everything itself — O1 and O3 have already proven that as long as reinforcement learning training is in place and context and memory are super long, ten thousand tokens is no problem, and the reasoning process is still clear. Multi-agent interaction definitely happens when a task involves an AI being passively triggered in the middle, for example when in the future every website's entrance is also an agent, or you have AI look something up on Doubao (豆包) or ChatGPT. But for the time being most human-computer interaction interfaces are still graphical interfaces, still designed for people, so the next year or two is still a single-agent state.
— Wu YiOperator's user data is worth far more than a chatbot's
Wu Yi agrees with Guangmi that a chatbot cannot help improve intelligence — small talk has no intelligence, and the dialogue form by default mostly won't produce a deliberative process. But Operator is different: when do you need a person to click webpages for you? Usually you have a purpose, like calculating personal income tax or booking a complex itinerary; the query distribution is very different from a chatbot's, relatively heavy. So if you can collect such user feedback and final completion signals, it helps improve the model's general capability. OpenAI has an option not to submit your user scripts; Wu Yi says this kind of instruction carrying a complex task, plus data on whether it was completed, is very well suited to reinforcement learning training, and if a hard problem is found, reinforcement learning training needs very little data.
— Wu YiOperator is a signal, not intelligence in the physical world
Wu Yi has long held a view: intelligence splits into two parts, one is reasoning in the pure text modality, the other is reasoning over visual signals, and the two are not quite the same. Operator now hasn't really gone into the physical world; it is still at the software webpage level, and webpages are mostly designed for people to browse, with a relatively strong textual concept. So what it shows is the beginning of a possibility; it hasn't really dragged the possibility out. It expanded capability, but the intelligence ceiling hasn't moved — pure logical reasoning is not stronger than O3, and physical-world reasoning hasn't arrived either. Wu Yi uses a pointer as a metaphor: originally it was a 90-degree pointer always heading north, and now it has expanded a bit to the right from 90 degrees, maybe to 80 degrees, and you don't know where it will go.
— Wu YiIn their own words · checked verbatim
It's that foundation model. If there's no good foundation model, reinforcement learning alone won't do. But look, a good foundation model alone won't quite do either, right? You still have to add reinforcement learning. The two together.
就是那个基础模型 如果没有好的基础模型 光靠强化学习是不行的 但是你看光靠好的基础模型 也不太行 对吧 还要加强化学习两块加起来
Wu Yi22:23
But innovation is different. Innovation is actually that you can only give something directional. For example, we know students, we know PhD students writing papers — you actually can't say whether this thing is right or wrong, because right or wrong has become meaningless. Because if it's innovation, it's definitely right.
但是创新不一样 创新其实是你只能给出方向性的 比如说我们知道学生 我们知道博士生写论文 你其实不能说这个东西是对的还是错的 因为对错没有意义了 因为如果要创新 它肯定是对的
Wu Yi25:28
So if through that kind of massive user feedback it can't improve intelligence, and can only improve the product's comfort, then that's fine. But I think Operator is relatively a little different.
所以如果通过用户那样大量的反馈 它不能提高智能 只能提高这个产品的舒适度 所以这个是没有问题的 但是我觉得operator 相对来说有一点点不一样
Wu Yi39:46
It only shows the beginning of this possibility. It hasn't really dragged this possibility out.
它只是展现了 这个可能性的开端 它没有把这个可能性 真正拖出去
Wu Yi1:00:05
Figures
| Number laid off from OpenAI's 2016 Web Agent team | twenty or thirty people | 21:23 |
| Operator's score on the OS World benchmark | 30-something | 53:57 |
| Wu Yi's time working at OpenAI | 2019 to 2020 | 2:04 |
Glossary
- CUA / Computer-Using Agent
- The model OpenAI custom-trained for Operator, specifically for interaction and agent use.
- GUI agent
- An agent that completes tasks by operating a graphical interface with mouse and keyboard.
- close loop system
- A system that receives environment feedback after outputting an action, then makes the next decision based on that feedback.
- OS World
- A benchmark evaluating an agent's ability to complete tasks in an operating system environment.
How to listen
Engineers, founders, and investors who want to understand the agent technical roadmap and Operator's positioning — especially those working on reinforcement learning, multimodality, or agent products.
The wrap-up and subscription outro after 1:11:21 can be skipped.