Companies that don't do reinforcement learning may not make it through the next wave
After language pretraining hit a bottleneck, the few hundred most core researchers in Silicon Valley have bet their resources on self-play RL, which trades inference compute for training compute and lets models explore logical reasoning on their own.
The video won't play here. Listen to the audio instead:
The argument · tap a timestamp to hear it
The pretraining path may have already hit a bottleneck
Guangmi's judgment: there is a 50% chance that the traditional Scaling Law has already failed, and the other 50% is that continuing down the old road and throwing 100,000 cards at it can still lead to AGI — the two probabilities are ‘Harvard Harvard’. The evidence he sees is that all three elements — parameters, data, and compute — are stuck. The best model is a 6700B-total-parameter MOE, because that is the limit of what H100s on a single server can hold, and no model has been seen pushing up three to five times to two or three trillion parameters. On data, many companies have scraped 15 to 20T of high-quality text, but it is hard to increase that to 50 to 100T. On compute, the largest single H100 cluster is 32,000 cards, and there will be no order-of-magnitude improvement before the B series comes out.
— Guang MiMultimodality and 100,000-card clusters are not paradigm-level changes
Guangmi groups the alternative routes into three. The first is multimodality, especially vision: high certainty, but today there is no evidence that intelligence or logical ability emerges from training on the visual modality — Tesla FSD doesn't work when moved to another new device, there is no generalization. The second is 100,000-card clusters, betting that compute decides life or death, but a 30,000-card cluster basically bricks once every two hours, and a 100,000-card cluster bricks once every twenty or thirty minutes; the interconnection difficulty may be more complex than SpaceX launching a heavy rocket. Both of these are a matter of time, but ‘not essential’. The only thing that truly counts as paradigm-level is reinforcement learning, RL.
— Guang MiRL lets AI try random paths by trial and error
Ilya, in a guest lecture at MIT in 2018, summed up reinforcement learning in one sentence: let AI try a new task along a random path, and if the result exceeds expectations, update the neural network's weights so that AI remembers to use this successful practice more, then start the next attempt. Guangmi uses a prospecting metaphor to explain the key to the reward model: one person holds a treasure map, another brings 5,000 special forces soldiers and professional detection equipment — as long as there is treasure, they can almost certainly find it, but if two or three of those special forces are not good at appraising treasure, they will miss the treasure or pick up garbage — that is the reward model making mistakes. Today the industry's reward models are still most core in code and math, because the environment and goals are simple and clear, easy to set up.
— Guang MiLanguage is a crutch; reinforcement learning is the main course
Guangmi offers a vivid metaphor: if language and pretraining are compared to the human genome, carrying the genes of thousands of years of human evolution, then reinforcement learning is the entire life of human growth, receiving positive and negative signals from the day of birth. Language is the only thing today that has achieved generalization; AlphaGo can only play Go, CV can only do face recognition, neither generalizes. So language and pretraining may just be a crutch, an intermediate dish, an appetizer, and the reinforcement learning that follows is the main course. He judges that today, relying on large language models alone may not reach AGI; AI may be a college student who is lopsided in Chinese, and to get a job, a new paradigm needs to be introduced.
— Guang MiRL trades inference compute for training compute
Guangmi points out that RL is not a model but a whole system, including agents, environments, actions, and reward models, and the two most important are the environment and the agent. Its approach in language models is essentially to trade inference time for training time, to solve the current situation where the model's marginal returns are temporarily diminishing as it scales up. He did the math: for models at the level of GPT-4 and Claude 3.5, synthesizing 1T of high-quality reasoning data costs about $600 million, and synthesizing 10T might cost $6 billion. But inference has relatively lower requirements for single-card performance and cluster scale; you don't necessarily have to use the top cards or 30,000-card or 100,000-card clusters — distributed clusters can also run RL inference.
— Guang MiConsensus exists only among a few hundred core researchers
The fact that an AI paradigm shift is happening has some consensus in Silicon Valley only among the most core researchers, possibly just a few hundred people, and has not fully spread. Many people know RL is important but don't know how to do it, and talent in this area is scarce — and it's not the traditional RL crowd. Guangmi thinks many AI executives may not yet be aware, because only a small number of papers have recently started to come out. Yann LeCun recently criticized reinforcement learning, saying it is a waste of resources; Guangmi's response: Edison also wasted a lot of experimental resources inventing the light bulb, but you only need to succeed once, and then you can replicate it en masse.
— Guang MiSilicon Valley invests in robot brains, China invests in whole machines
Guangmi observes that in Silicon Valley everyone now wants to invest in a technical person's brain, wanting to do ROS or Android; in China you invest in the whole machine — OV, Xiaomi. But he raises a paradox: is it possible that there is no such thing as a robot brain, that the brain is just GPT or a general large model, and that building a robot brain may not fit all hardware — data from machine A cannot be used on machine B, and end-to-end adaptation is still needed. He believes the most core, most core thing for general robots is still the timing of the technology; a general humanoid explosion may still be five to ten years away, and it is very likely this batch of companies won't truly make it — everyone is still at the stage of a research lab.
— Guang MiThe hidden line of mobile internet is recommendation; the hidden line of AI is RL
Guangmi compares mobile internet with today's AI: the visible line of mobile internet is that the world gained four to five billion mobile users, and the hidden line is having user behavior data to do recommendation — companies that didn't do recommendation over the past decade didn't grow big. The key feature capabilities were large screens, cameras, and GPS, each of which gave birth to very large companies. Today AI's visible line is still the Scaling Law, with Compute at its core; the hidden line may be Self-Play reinforcement learning. He raises a possibility: companies that don't do reinforcement learning today may not make it through the next wave, just like recommendation. The ranking of AI's key capabilities is Coding, multimodality, math, Agent.
— Guang MiIn their own words · checked verbatim
Right now we can only say there's a 50% chance that the traditional Scaling Law has already failed, and of course the other 50% chance is that following the old path can still lead to AGI, right? Keep throwing 100,000 cards at it. Feels like these two probabilities are Harvard Harvard.
现在只能说有50%的概率 就是传统意义上的Skilling Law已经失效了 当然另外50%的概率就是说 沿着老的路还能继续走向AGI对吧 继续怼十万卡 感觉这两个概率哈佛哈佛吧
Guang Mi3:08
That is, let AI try a new task along a random path; if the result exceeds expectations, update the neural network's weights, let AI remember to use this successful practice more, and then start the next attempt.
就是说让AI用随机的一个路径去尝试一个新的人物 如果效果超预期那就更新神经网络的权重 让AI记得多使用这个成功的时间 然后再开始下一次的尝试
Guang Mi13:20
You can compare language and pretraining to the human genome, carrying the genes of thousands of years of human evolution, then reinforcement learning is the entire life of human growth.
你可以把语言和预训训练 比作人类的一个基因组 携带着人类几千年进化的基因 那么强化学习就是人类成长的一生
Guang Mi23:31
Because the idea of RL is essentially to trade inference time for training time, to solve this problem of diminishing marginal returns.
因为R的思路 本质是用inference time 换training time 那来解决这个编辑收益递减的问题
Guang Mi34:45
Also wasted a lot of experimental resources, right? But you only need to succeed once, then you can replicate it en masse.
也浪费了大量的实验资源 对吧 但你只需要成功一次嘛 那你就可以大量复制
Guang Mi37:45
There's even a possibility that companies that don't do reinforcement learning today won't make it through the next wave, just like recommendation.
甚至说有没有一个可能性 今天不做强化学习的公司 下一波浪潮里面都跑不出来 这就跟推荐一样
Guang Mi1:07:02
In fact, every technological revolution goes through hardware investment first, then infrastructure building, then an application explosion. Historically, too, there was railway construction first, then economic activity later.
其实每一次科技变革都是经历先硬件投入 再infra建设再应用爆发 历史上也都是先有铁路建设 再有后来的经济活动
Guang Mi1:18:11
Figures
| Maximum interconnection scale of a single H100 cluster | 32,000 cards | 4:09 |
| High-quality text data volume | 15 to 20T | 4:09 |
| Brick frequency of a 30,000-card cluster | once every two hours | 8:15 |
| Brick frequency of a 100,000-card cluster | once every twenty or thirty minutes | 8:15 |
| Cost of synthesizing 1T of high-quality reasoning data | about $600 million | 34:45 |
| Cost of synthesizing 10T of high-quality reasoning data | about $6 billion | 34:45 |
| Character.AI acquisition price | over $2 billion | 31:39 |
| OpenAI annualized revenue | $4 billion, possibly $7-8 billion by year-end | 1:12:06 |
Glossary
- Self-Play RL
- A new paradigm in which a model repeatedly plays against itself or the environment to explore, trading inference compute for training compute.
- Scaling Law
- The empirical rule that model capability improves predictably as parameters, data, and compute increase.
- reward model
- In reinforcement learning, the model that judges whether an AI action is good or bad and gives positive or negative signals.
- RLHF
- Using human preferences to train a reward model, with the goal of making AI more human-like and achieving human-machine alignment.
- MOE
- A model architecture composed of multiple expert sub-networks, activating only part of the parameters at a time.
- DiT
- The video generation route indicated by Sora, using a Transformer to do diffusion models.
How to listen
Founders, investors and engineers watching the evolution of AGI technical routes, especially those trying to judge where the next wave of opportunity lies after pretraining.
After 1:05:54, the part about basic research culture and Silicon Valley player commentary can be fast-forwarded.