The world is too loud. Read what matters.

张小珺·商业访谈录

o1 Is Not a Detour but an Advance: Fast-Thinking Models Learned to Think Slowly

Wang Xiaochuan argues o1 is a paradigm upgrade, not two steps back for one forward — you need fast thinking first before reinforcement learning can teach it to think slowly, and this path has touched the beginnings of the route from language to intelligence.

Reinforcement Learningo1Medical AIStartup StrategyLarge Models
Wang Xiaochuan returns, and lays out o1's reinforcement learning principles, the difficulty of reproducing it, medical deployment and the range of a startup more concretely than most analyses — suited to anyone who wants the industry view.

The argument · tap a timestamp to hear it

5:45

o1 is not a detour, it is a paradigm upgrade

Wang Xiaochuan does not accept the ‘two steps back for one forward’ framing. He sees o1 as a paradigm upgrade: you first need fast thinking before you can use a large model's fast thinking to teach it slow thinking — that is an advance, not a detour. He uses the DIKW model to explain: search sits at the Information layer, large models LLM reached the Knowledge layer, and o1's slow thinking lets the model begin evolving from Knowledge toward the beginnings of Wisdom — it starts to have intelligence. The original model becomes one component of o1, rather than reinforcement learning merely serving the large model.

— Wang Xiaochuan
9:04

o1 hiding its chain of thought is a competitive strategy

o1 hides its thinking process, and cracking the chain of thought gets you a warning and a ban. Wang Xiaochuan gives two reasons: first, previously everyone could rapidly approach the frontier by distilling from large-model data, and OpenAI is a commercial company, not a charity — once it goes public it is easier for others to imitate the logic and distill the data, letting others progress faster; second, it shows the unique information in this technology itself is limited. So the lockdown is a competitive strategy. He also points out o1's two core focuses: sticking to language as the center and moving toward CoT, and splitting the thinking process and the result into two stages, which lets CoT generalize better.

— Wang Xiaochuan
11:38

Reinforcement learning does not tell the process, only judges right or wrong

Wang Xiaochuan uses teaching a child as an analogy: supervised learning has to tell the solving process, step one, two, three, and the child learns fast but does not know why; reinforcement learning does not tell the process, it only judges whether the work is right or wrong — right is called right, wrong is called wrong — and the child goes in and finds the method itself. He explains why large models especially need reinforcement learning: a large model trains and compresses the finest language in the world, it is intelligence within the original data distribution, and its thinking ability cannot exceed the original data; but talking about intelligence means jumping outside the framework and moving out of distribution, so you must create an environment and use environmental feedback to bring in content beyond language data.

— Wang Xiaochuan
16:39

Medicine is a good domain for reinforcement learning

Wang Xiaochuan thinks doctors are a rather good domain to improve with reinforcement learning, because many medical questions have standard answers: what symptoms the patient has, what tests and exams should be done, what medicine should be prescribed — all have answers. If you access the doctor's CoT and then verify whether the answer is right, the model's power could rise sharply. A doctor does not learn just by reading medical school books; over a clinical career they may see tens of thousands of patients, improving through interaction with patients, and much of that data gets recorded. He also mentions Baichuan earlier ran experiments with Tang and Song poetry, because the cipai have rules for character count, tonal patterns, rhyme and parallelism, so a program can serve as a reward model to check them.

— Wang Xiaochuan
19:55

Earlier reinforcement learning had no CoT

Wang Xiaochuan admits Baichuan's earlier reward model did not carry CoT, and after seeing o1 today they want to reproduce a stronger CoT. Having CoT has two meanings: first, in medicine it lets doctors give their thinking path earlier and improve capability faster, not just end-to-end; second, generalization ability rises substantially — if the line of thinking is right, the answer is right. He also shares an observation: reinforcement learning partly learns new things from the environment, and partly activates existing abilities. They taught the model character count, tonal patterns and rhyme, and the model on its own also output parallelism that had not been taught, showing that latent memory and ability can be activated.

— Wang Xiaochuan
22:16

Reproducing o1 may be faster than reproducing GPT-4

Wang Xiaochuan thinks reproducing o1 will be somewhat faster than reproducing GPT-4 was back then. Hard as it is, China and the US have a large number of open-source projects, big companies and startups are entering, and capital abundance and talent concentration are already far greater than after the GPT-3.5 and GPT-4 releases. He expects a model close to o1 to appear within one or two months, though reaching market height will still take effort. He estimates GPT-4 may have taken 18 months, o1 perhaps 9 months, and from a standing start something may appear in one or two months. He also says o1 is like the GPT-3 release back then, but because it runs a new paradigm through on top of GPT-4, its importance is no less than GPT-3's.

— Wang Xiaochuan
26:30

Code will play a more important role in the future

Wang Xiaochuan predicts code will play a more important role in the future: before, code helped improve logical ability and assisted programmers in writing code; in the future code will become the large model's next core capability, with the large model solving more problems by writing code, even solving its own thinking process. He believes reinforcement learning will also move toward a new paradigm of writing code to solve problems, to be realized in the coming years. He also mentions o1's visible ceiling: perhaps within two or three years the paradigm will run out an interface, and the rest is code playing a more important role — continuing to write its own code, using code to run and generate neural networks, even fusing neural networks and models together, and a new paradigm will emerge.

— Wang Xiaochuan
28:35

Build applications that rise with the tide, not eggs laid along the way

Wang Xiaochuan divides applications into two kinds: laying eggs along the way means whenever the model gets better I lay a new egg — first an advertising model, then a customer-service model — and the more eggs, the more pressure; rising with the tide means the bigger the model, the better this domain does, rather than a stage where a big model has nothing to do with my domain. What he looks for is the application scenario in its ultimate form — assuming model capability becomes especially strong, which scenario benefits most, while also being enterable when model capability is ordinary, with a threshold that is not that high, but benefiting more the bigger the model gets. This year he began boldly raising medicine; next year it is a two-wheel-drive model, and he hopes to get tested by the market.

— Wang Xiaochuan
31:35

Startups must get out of the big companies' range

Talking about the six little dragons, Wang Xiaochuan says at least one will survive, and says he is both protection and attack, and in areas where there is also consensus he will develop very fast. He believes there must be something higher than the big companies that they cannot see, or that their organizational capability cannot do, for a startup to have a chance to survive — this is called getting out of the big companies' range; inside the range you have no good way to live. He also redefines Baichuan: a first-tier, clear-thinking large-model company, the only one. Asked about the hardest moment, he says it was building the team at the start; once the team was there, things got better.

— Wang Xiaochuan

In their own words · checked verbatim

That also shows the unique information in this technology itself is limited, so its locking things down is a competitive strategy.

那也说明这个技术本身 它的这个独有信息也有限的 所以它封锁这个事情 是一个竞争策略

Wang Xiaochuan9:04

A doctor writes because he does not learn just by reading medical school books and being done; in clinical practice he may see tens of thousands of patients in a lifetime and has to improve himself, so a doctor improves through interaction with patients.

医生写是因为他 不是光看医学院的书 读完了就会了 他在临床中间 大概一辈子可能看几万个病人 要得到自己的提升 所以医生是靠病人的这个 这个互动当中去得到自己的提升的

Wang Xiaochuan18:09

It is that the bigger the model, the better my domain can do — rather than a stage where a big model has nothing to do with my domain.

就是模型越大 那我这个领域能做得更好 而不是模型大的一个阶段 就跟我领域没关系了

Wang Xiaochuan28:35

Figures

Estimated time to reproduce GPT-418 months23:15
Estimated time to reproduce o19 months23:15
Time for a model close to o1 to appearone or two months23:15
Number of Chinese cipai100-plus19:10
Number of patients a doctor sees in a lifetimetens of thousands18:09

Glossary

CoT / chain of thought
Having the model write out its reasoning step by step rather than giving the answer directly.
reward model
In reinforcement learning, the evaluation system that judges whether an output is right or wrong and provides the training signal.
DIKW / Data-Information-Knowledge-Wisdom model
A four-layer progression framework from Data to Information to Knowledge to Wisdom.
Self Play RL
The model generates training data by competing against itself, reducing manual labeling.
scaling law
The empirical rule that model capability grows with compute, data and parameter scale.

How to listen

Who it's for

Founders, investors and engineers watching large-model technical routes and startup strategy, especially those who want to understand o1's reinforcement learning principles and the path to medical deployment.

Skip

The show intro and the Sam Altman palace-intrigue setup at 0:00-1:00 can be fast-forwarded.