The world is too loud. Read what matters.

张小珺·商业访谈录

o1 Is Not a Change of Course — It Just Added a Few More Rungs to the AGI Ladder

The pretraining gold mine is nearly exhausted; reinforcement learning is the second mine. o1's preview looks more like a GPT-3 moment — whether it becomes a ChatGPT moment depends on the official release.

Reinforcement Learningo1Post-trainingAGIOpenAI
A former OpenAI researcher explains the three elements of reinforcement learning behind o1, why reward models are so hard, and how OpenAI used to mine with its eyes closed.

The argument · tap a timestamp to hear it

7:20

The pretraining gold mine is nearly dug out

Wu Yi compares AGI to mining and building ladders: pretraining is the first big gold mine, and after all these years of digging, there is less and less left to extract. Post-training is the second big gold mine, where you can do reinforcement learning, exploration, search and generate synthetic data — and it will very likely feed back into pretraining in turn. So o1 is not a change of course; it is that ‘stage one is over, the pure-pretraining stage is over’, and we have entered a post-training stage built on reinforcement learning, adding a few more rungs to the ladder toward AGI.

— Wu Yi
9:00

o1's preview looks more like a GPT-3 moment

Wu Yi thinks a direct analogy is hard, but he leans toward the preview version possibly being a GPT-3 moment, with the ChatGPT moment waiting for the official release. The reason: ChatGPT and GPT-3 had no essential difference in base capability — RLHF and instruction following were just done better, making the model usable and productized, which is why it suddenly took off. o1 has already moved up a notch in capability, but whether it can make people who previously found it unusable find it usable as a product is still unknown.

— Wu Yi
15:06

Every one of the three elements of RL is hard

Wu Yi breaks reinforcement learning into three parts: reward model, search and exploration, and prompt. He uses the analogy of coaching a middle-schooler for competitions: the teacher giving feedback is the reward model, what difficulty of problems to set is the prompt, and whether the student can generalize and figure out on their own how to get it right is exploration. All three matter, and they must all be done right at the same time to achieve capability gains. A good reward model alone cannot guarantee capability improvement — this is why reinforcement learning has such a high barrier.

— Wu Yi
19:36

PPO is only useful when you have enough compute

DQN learns a value network and infers the Agent from it; the optimization is indirect and has a gap, but its mathematical properties are good and its demands on compute and data are small. PPO trains the policy network directly, letting the policy explore and evolve on its own — more intuitive, but with worse mathematical properties, and it needs enough exploration, so ‘PPO is useful if and only if you have enough compute’. Without enough compute, running PPO has no effect at all, which is also why academia more often uses Offline RL algorithms like DQN and DPO.

— Wu Yi
33:24

A reward model cannot be trained alone with eyes closed

Wu Yi draws an analogy to P and NP in theoretical computer science: the reward model is an NP-type problem — judging whether an answer is good is indeed simpler than writing the answer, but not by much. And many problems may have no universal reward model, because human preferences have no consensus. His advice: it is unlikely that a separate team can train a reward model with its eyes closed; it must be coupled — progress in the model's reasoning ability produces higher-quality data, which brings a better reward model, and a better reward model in turn drives the model forward.

— Wu Yi
40:31

Hallucination comes from models knowing correlation, not causation

Wu Yi sees two causes of hallucination. First, the model does not know causality, only correlation, so it does not know whether it actually knows the answer — ask it who won the World Cup and it says Brazil. Pretraining and SFT have no counterfactual reasoning process and easily overfit to correlations. Second, many reasoning problems have a great many intermediate steps, and the traditional AI paradigm demands outputting the correct answer in one shot, with no allowance for revision. o1 gives the model a buffer of 10 or 20 seconds, allowing it to explore and to revise — and that alone can greatly improve reasoning performance.

— Wu Yi
49:49

o1 does not mean vertical models will rise

Wu Yi firmly believes a general, unified model will emerge. The reason: when the parameter count is very large, many vertical models are easily orthogonal in high-dimensional space, and orthogonal parameters are easily merged. But he stresses the premise is that the vertical method is coupled with the pretraining paradigm — a customized mathematical method like Alpha Geometry cannot feed back into a large model's base capability, whereas if OpenAI is doing domain-specific training on general methods, then a better vertical model must mean a better general model. He expects the RL paradigm to become common in 2 to 4 years, but within a 2-year horizon there will not be very many vertical models.

— Wu Yi
1:03:20

OpenAI's KPI back then was blog readership

Wu Yi describes early OpenAI as a bizarre organization that was ‘product-driven yet had no product’: each group's KPI was to make a big splash, and a big splash meant publishing a blog post. The multi-agent team worked on hide and seek for over a year and published one blog post, and that blog post may have been worth tens of millions of dollars. He thinks this model fit the Scaling Law path — if you bet on Scaling Law, a project cannot be built by one or two excellent researchers plus one or two helpers; it requires heavy engineering investment, and the paper cycle is too short to serve as an evaluation. This model was not actually invented by OpenAI either — DeepMind did it this way when it built AlphaGo.

— Wu Yi

In their own words · checked verbatim

It is unlikely that there could be a separate team that, regardless of the model's own capability, trains a reward model with its eyes closed — that is probably not possible.

reward model这个东西不太有可能说有一个单独的小组,我不管这个模型的本身的能力,我闭着眼睛去训练一个reward model,这件事情恐怕是不太可能的。

Wu Yi36:28

The traditional AI chain of thought, or this output mode, does not allow you to revise — this is actually one cause of hallucination. Often the model could revise to the right answer, but you simply never gave it the chance to revise.

传统的AI的这样思维链也好,还是这样输出的模式也好,它是不允许你改的,这个其实也是幻觉的一个原因,很多时候可能这个模型可以改对,但你根本没有给它改的机会。

Wu Yi43:33

Figures

o1 reasoning chain lengthseveral thousand tokens5:04
o1 reasoning timeabout 10 seconds5:04
Wu Yi's time at OpenAIFebruary 2019 to the end of July 20202:00
OpenAI headcount inflection pointunder 100 people before 20211:00:45
Number of people on Wu Yi's team doing algorithmsfive or six people54:41

Glossary

PPO / Proximal Policy Optimization
A reinforcement learning algorithm proposed by OpenAI in 2017 that trains the policy network directly and is only effective with enough compute.
DQN / Deep Q-Network
An algorithm proposed by DeepMind in 2014 that trains a value network and then infers the Agent, with low compute requirements.
RLHF / Reinforcement Learning from Human Feedback
A method that uses human feedback to train a reward model and then align large-model behavior.
self-play
Having a model play against itself to improve, suited to problems with symmetric structure such as Go.
Scaling Law
The regularity that increasing compute, data and parameter scale brings improvements in model capability.

How to listen

Who it's for

Engineers, investors and founders watching large-model technical directions, reinforcement learning and AI startup opportunities — especially those who want to understand o1's underlying mechanism rather than just look at benchmark scores.

Skip

The first 3 minutes of guest self-introduction and the idle chat about choosing between two offers can be skipped.