o1 Is Not a Change of Course — It Just Added a Few More Rungs to the AGI Ladder
The pretraining gold mine is nearly exhausted; reinforcement learning is the second mine. o1's preview looks more like a GPT-3 moment — whether it becomes a ChatGPT moment depends on the official release.
The video won't play here. Listen to the audio instead:
The argument · tap a timestamp to hear it
The pretraining gold mine is nearly dug out
Wu Yi compares AGI to mining and building ladders: pretraining is the first big gold mine, and after all these years of digging, there is less and less left to extract. Post-training is the second big gold mine, where you can do reinforcement learning, exploration, search and generate synthetic data — and it will very likely feed back into pretraining in turn. So o1 is not a change of course; it is that ‘stage one is over, the pure-pretraining stage is over’, and we have entered a post-training stage built on reinforcement learning, adding a few more rungs to the ladder toward AGI.
— Wu Yio1's preview looks more like a GPT-3 moment
Wu Yi thinks a direct analogy is hard, but he leans toward the preview version possibly being a GPT-3 moment, with the ChatGPT moment waiting for the official release. The reason: ChatGPT and GPT-3 had no essential difference in base capability — RLHF and instruction following were just done better, making the model usable and productized, which is why it suddenly took off. o1 has already moved up a notch in capability, but whether it can make people who previously found it unusable find it usable as a product is still unknown.
— Wu YiEvery one of the three elements of RL is hard
Wu Yi breaks reinforcement learning into three parts: reward model, search and exploration, and prompt. He uses the analogy of coaching a middle-schooler for competitions: the teacher giving feedback is the reward model, what difficulty of problems to set is the prompt, and whether the student can generalize and figure out on their own how to get it right is exploration. All three matter, and they must all be done right at the same time to achieve capability gains. A good reward model alone cannot guarantee capability improvement — this is why reinforcement learning has such a high barrier.
— Wu YiPPO is only useful when you have enough compute
DQN learns a value network and infers the Agent from it; the optimization is indirect and has a gap, but its mathematical properties are good and its demands on compute and data are small. PPO trains the policy network directly, letting the policy explore and evolve on its own — more intuitive, but with worse mathematical properties, and it needs enough exploration, so ‘PPO is useful if and only if you have enough compute’. Without enough compute, running PPO has no effect at all, which is also why academia more often uses Offline RL algorithms like DQN and DPO.
— Wu YiA reward model cannot be trained alone with eyes closed
Wu Yi draws an analogy to P and NP in theoretical computer science: the reward model is an NP-type problem — judging whether an answer is good is indeed simpler than writing the answer, but not by much. And many problems may have no universal reward model, because human preferences have no consensus. His advice: it is unlikely that a separate team can train a reward model with its eyes closed; it must be coupled — progress in the model's reasoning ability produces higher-quality data, which brings a better reward model, and a better reward model in turn drives the model forward.
— Wu YiHallucination comes from models knowing correlation, not causation
Wu Yi sees two causes of hallucination. First, the model does not know causality, only correlation, so it does not know whether it actually knows the answer — ask it who won the World Cup and it says Brazil. Pretraining and SFT have no counterfactual reasoning process and easily overfit to correlations. Second, many reasoning problems have a great many intermediate steps, and the traditional AI paradigm demands outputting the correct answer in one shot, with no allowance for revision. o1 gives the model a buffer of 10 or 20 seconds, allowing it to explore and to revise — and that alone can greatly improve reasoning performance.
— Wu Yio1 does not mean vertical models will rise
Wu Yi firmly believes a general, unified model will emerge. The reason: when the parameter count is very large, many vertical models are easily orthogonal in high-dimensional space, and orthogonal parameters are easily merged. But he stresses the premise is that the vertical method is coupled with the pretraining paradigm — a customized mathematical method like Alpha Geometry cannot feed back into a large model's base capability, whereas if OpenAI is doing domain-specific training on general methods, then a better vertical model must mean a better general model. He expects the RL paradigm to become common in 2 to 4 years, but within a 2-year horizon there will not be very many vertical models.
— Wu YiOpenAI's KPI back then was blog readership
Wu Yi describes early OpenAI as a bizarre organization that was ‘product-driven yet had no product’: each group's KPI was to make a big splash, and a big splash meant publishing a blog post. The multi-agent team worked on hide and seek for over a year and published one blog post, and that blog post may have been worth tens of millions of dollars. He thinks this model fit the Scaling Law path — if you bet on Scaling Law, a project cannot be built by one or two excellent researchers plus one or two helpers; it requires heavy engineering investment, and the paper cycle is too short to serve as an evaluation. This model was not actually invented by OpenAI either — DeepMind did it this way when it built AlphaGo.
— Wu YiIn their own words · checked verbatim
It is unlikely that there could be a separate team that, regardless of the model's own capability, trains a reward model with its eyes closed — that is probably not possible.
reward model这个东西不太有可能说有一个单独的小组,我不管这个模型的本身的能力,我闭着眼睛去训练一个reward model,这件事情恐怕是不太可能的。
Wu Yi36:28
The traditional AI chain of thought, or this output mode, does not allow you to revise — this is actually one cause of hallucination. Often the model could revise to the right answer, but you simply never gave it the chance to revise.
传统的AI的这样思维链也好,还是这样输出的模式也好,它是不允许你改的,这个其实也是幻觉的一个原因,很多时候可能这个模型可以改对,但你根本没有给它改的机会。
Wu Yi43:33
Figures
| o1 reasoning chain length | several thousand tokens | 5:04 |
| o1 reasoning time | about 10 seconds | 5:04 |
| Wu Yi's time at OpenAI | February 2019 to the end of July 2020 | 2:00 |
| OpenAI headcount inflection point | under 100 people before 2021 | 1:00:45 |
| Number of people on Wu Yi's team doing algorithms | five or six people | 54:41 |
Glossary
- PPO / Proximal Policy Optimization
- A reinforcement learning algorithm proposed by OpenAI in 2017 that trains the policy network directly and is only effective with enough compute.
- DQN / Deep Q-Network
- An algorithm proposed by DeepMind in 2014 that trains a value network and then infers the Agent, with low compute requirements.
- RLHF / Reinforcement Learning from Human Feedback
- A method that uses human feedback to train a reward model and then align large-model behavior.
- self-play
- Having a model play against itself to improve, suited to problems with symmetric structure such as Go.
- Scaling Law
- The regularity that increasing compute, data and parameter scale brings improvements in model capability.
How to listen
Engineers, investors and founders watching large-model technical directions, reinforcement learning and AI startup opportunities — especially those who want to understand o1's underlying mechanism rather than just look at benchmark scores.
The first 3 minutes of guest self-introduction and the idle chat about choosing between two offers can be skipped.