AI Is the Wave, the Surfer Doesn't Matter: Confessions of a Researcher at a Model Giant
Pretraining hasn't hit a wall — most people who think it has have a bug in their own work. AI is fundamentally simple, because it can run experiments. The era of individual heroism is over; what matters now is organizational systematicity and reliability.
The video won't play here. Listen to the audio instead:
The argument · tap a timestamp to hear it
Nobody worries about catching up now, only about what to do
Yao Shunyu says that a year ago at Anthropic, what everyone worried about was whether OpenAI's reasoning was so strong that they had a chance to catch up; now, among Gemini, OpenAI and Anthropic, none of the three would really worry about not being able to catch up. The harder thing has become figuring out what to do — that's a bet, and it really needs human insight. This is a different matter from ‘model capabilities being flattened’: in terms of actual user experience the three still differ, but on paper (benchmark) they're all around 80%, and being a bit higher or lower is mainly noise rather than signal.
— Yao ShunyuMost people who hit the pretraining wall have a bug in their own work
Many people were discussing months ago whether the pretraining scaling law had run its course; Yao Shunyu's experience is that it hasn't, and he sees no sign of it ending in the next four months either. He gives three possible reasons for ‘thinking it's over’: one, believing the law's scope of applicability has ended; two, believing some condition (like data hitting a wall) can't be met; three — and his impression is that the vast majority of people fall into this category — there's a bug in your own work that you haven't found. The bug could be a scientific hypothesis done wrong (for example, token horizon, or how much data to allocate per model size wasn't chosen clearly), or it could be a pure code bug. He stresses that the progress from fixing one bug is often far greater than from some very magical trick.
— Yao ShunyuCoding exploded because the reward signal is easy to define
Coding has been developing at high speed since the Claude 3.5 (called 3.6 externally) generation, and Yao Shunyu gives two model-level reasons: one, the reward signal is very easy to define — write a feature, certain inputs produce certain outputs, and you can test whether it's right immediately; two, the data has a natural foundation — GitHub has accumulated decades of code from high-quality programmers, and from that code you can build a great many environments. At the product level there's a third reason: good programmers write code in a fairly similar style — concise, clean, with reasonable abstractions — which makes Coding products much simpler than social or gaming products, where everyone's taste differs and you have to rely on recommendation algorithms.
— Yao ShunyuHard distillation means you don't know what you want to do; smart distillation is true multi-agent
Dario named three companies that distill from him. Yao Shunyu distinguishes two kinds: hard distillation is taking a bunch of generated tokens from Claude and forcibly training on them — commercially unethical and intellectually stupid, because the companies doing this fundamentally don't know what they want to do, and the only thing they can do is copy others to make their data look good. But distillation also has interesting scientific questions: if in your own data-generation chain you use other models as assistants, or use your own model to generate answers and other models as evaluators, that's a gray area but technically very interesting — different companies' models have very different language distributions, and merging them into one training system is true multi-agent, and Chinese labs may thereby become pioneers in multi-agent training.
— Yao ShunyuRobots haven't reached their GPT-1 moment yet
Yao Shunyu searched Amazon for the prices of Chinese robots and was surprised at how cheap they were, which he sees as reflecting China's advantage in the hardware supply chain. But on the software side he doesn't quite see it: robot models are currently at the stage of ‘optimizing for a given environment, a given scenario’ — doing RL and building virtual environments can improve things, but there's no strong generalization. He sees generalization as the watershed for many AI directions — over a decade ago you could train very strong models for translation or semantic analysis, but you couldn't improve all capabilities horizontally; language crossed into that stage after Transformer and GPT, and robots haven't yet. He recommends everyone go look at robotics labs — they're much more interesting than language model labs.
— Yao ShunyuAnthropic can be top-down, OpenAI can't
Yao Shunyu thinks it's very unique that Anthropic can implement a fairly top-down mechanism. The difficulty is: the person making technical decisions must also be a company decision-maker — technically they must command respect, so researchers will trust them; and they must be a company decision-maker, so they can take responsibility for the company. Anthropic has this condition: technical leaders like Jared Kaplan and Sam are company cofounders, they make the decisions themselves, it's their company. Other model companies find it hard: OpenAI can't do it, and Gemini finds it relatively hard too. He adds that big companies and startups naturally play differently — what matters for a startup is making a bet, you have to gamble on one thing, and top-down has an organizational advantage over OpenAI.
— Yao ShunyuAI is fundamentally simple, because it can run experiments
Yao Shunyu says this isn't a conclusion, it's a statement of his, which could be right or wrong. He explains: the point where AI is fundamentally simple is that it can run experiments. The difference from physics is that in physics, without experimental data at that energy scale, you simply cannot understand the theory at that scale; but AI isn't bound by this — if you can't understand it, that's fine, it can still move forward. He can in fact now run any experiment he can think of, it just takes time to scale up compute and prepare infrastructure, there's no fundamental difficulty. So AI doesn't give people the feeling of hitting a wall: there are lots of things to try, and it's not that you're empty-headed with nothing to try — there are too many ideas and you have to try them one by one.
— Yao ShunyuPretraining is also a kind of RL; the difference is in the data distribution
Yao Shunyu says it's hard to say from a purely technical angle what the essential difference is between pretraining, supervised learning and RL, because pretraining and SFT are essentially no different — it's just treating the data you get as ground truth, as expert output, and moving toward that distribution. Reinforcement learning is broader: what it outputs isn't a given expert, it's self-generated, containing both good and bad results, and you move toward the good and pull away from the bad. So pretraining and SFT are a subset of reinforcement learning. But these two things really are different in this era, and the biggest difference is in the data: pretraining data needs a good enough distribution, broad enough coverage, and quality doesn't need to be very high; post-training is the reverse — the distribution is far narrower, but the data you do have must be of very high quality.
— Yao ShunyuIn their own words · checked verbatim
I think the AI direction is fundamentally simple — there's no... I think there's no... except maybe the jump, that idea might require some very deep insight, but in the process afterward, many ideas are actually very trivial! Very stupid — anyone can think of them, anyone can do them, it's just that you anticipate well and hit the opportunity to do it.
我觉得AI这个方向本质上是简单 就是没有哪 我觉得没有哪个 除了可能跳变那一下 那个idea可能是得有一些很深刻的洞见 在之后的那个过程中 很多想法其实是非常trivial !就是非常愚蠢的 就是谁都能想 谁都能干 只是你预计好 撞到这个机会去干而已
Yao Shunyu2:32:13
But I think in the next six to twelve months, what it currently can't do is whether it can, from start to finish, complete an AI research task — for example, not only write the code, but also run the experiment, and after running the experiment see the result, and after seeing the result analyze the result, and after analyzing the result guide where it went wrong, then propose a new hypothesis, design new code, run a new experiment. This chain isn't complete yet, but I think this chain is probably what will gradually become complete next.
但是我觉得未来六到十二 他目前还做不到的事情是什么 是说他能不能从头到尾的 把一件AI研究的事做完 比如说他不仅能写这个code 他还能跑这个实验 跑这个实验还能看到这个结果 看到这个结果 还能分析这个结果 分析这个结果 指导他哪做的不对 然后提出新的假设 设计新的代码 跑新的实验 这条链条目前还没有完整 但我觉得这条链条 可能是下一步会慢慢变得完整的事
Yao Shunyu2:38:17
I think everyone now is a surfer — essentially it's a wave, not you the surfer. Is the wave AI? Yes, the AI thing itself is the wave, it will move forward, whether you surf it or not, the wave will hit the shore, it's just that some people may surf it, and some may be a bit late and miss the crest.
我觉得大家现在就是 每个人都是冲浪的人 本质上是一个浪 而不是你那个冲浪的人 浪是AI吗 对 就是AI这个事情本身是这个浪 它会往前走 不管你冲不冲这个浪 这个浪都会拍到岸上 只是说有人可能就冲了这个浪 有人就可能晚了一点 没赶上这个浪尖
Yao Shunyu3:06:42
I think if a researcher can't consider the whole picture, then he isn't a qualified researcher in this era — and I think this is very different from doing research in academia, because doing research in academia is essentially a state of one person fed and the whole family not worried, I'm responsible for my project, right, I'm responsible for my reproducibility. But in a company, more often you have to be responsible for the company. These are two completely different mindsets.
我觉得如果一个研究员做不到对全局去考虑的话 他就不是考的研究员 在现在这个时代 就是这个和我觉得这个和你就是在学术界做research是很不一样的事 因为在学术界做research本质上是一个人吃饱全家不愁的状态 就是我为我的项目负责 对吧 我为我的可重复性负责 但是在一个公司里你其实更多的时候是我得为这个公司负责 对这是两种完全不一样的心态
Yao Shunyu3:18:56
Figures
| Anthropic employee count (when Yao Shunyu joined) | seven to eight hundred | 1:58:36 |
| Anthropic employee count (when Yao Shunyu left) | close to 2000 | 2:22:05 |
| Size of Yao Shunyu's large team at Anthropic (when he joined) | around 10 | 1:58:36 |
| Time from start of training to release for Claude 3.7 | four to five months | 2:18:04 |
| Yao Shunyu's estimate of the share of code generated by models | conservatively over 90%, non-conservatively maybe 99 or 100% | 39:29 |
| Yao Shunyu's estimate of experiment efficiency improvement | 20 to 50 times faster than a year to a year and a half ago | 41:30 |
| Gemini's chatbot market share (Yao Shunyu's estimate) | around 20% | 2:52:34 |
| Actual time Yao Shunyu spent as a Berkeley postdoc | two weeks | 1:43:30 |
| Yao Shunyu's view on when AI will run its own experiments | the next six to twelve months | 2:37:16 |
| Proportion Yao Shunyu publicly cited for the visit-China factor in his departure | 40% | 2:26:07 |
Glossary
- scaling law
- An empirical rule describing how model performance changes with model size and data volume.
- reward signal
- The signal in reinforcement learning used to judge whether a model's output is good or bad; easy to define in Coding scenarios.
- multi-agent
- Multiple models collaborating or acting as each other's evaluators; when different companies' models have different language distributions, it's closer to true multi-agent.
- VLA
- The model direction that uses a language model as the base and then trains robot actions.
- context management
- Selectively discarding and retrieving information under a limited context window to complete longer tasks.
- continual learning
- A learning approach that lets a model continuously update its weights during operation without forgetting old knowledge.
How to listen
AI practitioners, investors and founders watching model-company organizational mechanisms and pretraining vs. post-training roadmap judgments; anyone wanting to understand the internal differences between Anthropic and Google DeepMind.
1:20:59 to 1:53:47, the non-Hermitian systems, high-energy physics and physics-to-AI section, unless you're interested in the path of a researcher with a physics background.