Pretraining Is the Main Peak: Intelligence Itself Is the Biggest Application
Model companies are refineries; the value is in chemical plants and car companies. But today's biggest dividend is capturing intelligence spilling out of research, not building product-pulled applications.
The video won't play here. Listen to the audio instead:
The argument · tap a timestamp to hear it
Pretraining is over is the biggest consensus, and also the biggest non-consensus
Guangmi believes the biggest consensus today is that "Pretraining is over," but the biggest non-consensus is precisely that "Pretraining still has a huge amount of room, and has even just begun." His judgment: only Pretraining can produce emergent new capabilities; Post Training and RL merely elicit or strengthen capabilities, and determine the model's intrinsic ceiling. He uses an analogy: doing reinforcement learning on a very poor Base Model is like an elementary school student grinding practice problems—it saturates and hits a ceiling easily; only Pretraining can essentially turn an elementary school student into a middle school student. He judges that the next generation of SOTA models will still significantly beat today's SOTA.
— Guang MiOpenAI not prioritizing pretraining is an organizational problem, not a strategic one
From the outside, OpenAI seems to no longer prioritize Pretraining. Guangmi gives two reasons: first, strategic choice—the O series moves fast on benchmarks, two months of score-grinding may yield returns faster than one or two years of Pretraining, and ChatGPT's growth is also terrifying, consuming a large share of management's energy; second, organizational problems—OpenAI's core Pretraining team has been in constant turmoil, with Daryl from Anthropic taking away a batch early on, Elia leaving, and CTO Mira taking away the core Post Training people, while the original Pretraining people were continually reassigned to Post Training. So it's not that top-down doesn't value it, but that organizational adjustments make it appear less valued.
— Guang MiCoding is the model's hand, and also the best environment for AGI
Guangmi says he figured one thing out: Coding may be the best environment for achieving AGI. The significance of Coding is not in programming itself, but in that the vast majority of tasks in the real world can be expressed through code. He analogizes Coding to AlphaGo's board, Baidu's general search, and Taobao's product search environment. Coding is a hand of the model—by generating and executing code, the model achieves collection, processing, and feedback with the external environment. He judges that within two years, Agents will be able to operate digital environments like computers and phones, with capabilities absolutely surpassing 99% of people, performing 99% of normal human actions, and doing it better than humans.
— Guang MiOpenAI and Anthropic share the same origin, but their strategies have diverged
Guangmi believes OpenAI and Anthropic started with the same route, but as they went on, their core strategic bets changed greatly. OpenAI's core bets are two: first, hoping to achieve AGI through the O series or Reasoning Models; second, hoping to make ChatGPT into a K-Lab with over a billion active users to overturn Google first. Anthropic's core bet is to focus on Pretraining, continue to bet on Pretraining, have a very strong Base Model, and continue to bet on Coding and Agentic. OpenAI places more emphasis on the consumer market, Anthropic on the enterprise market; OpenAI has a bottom-up innovation organizational culture, Anthropic a top-down one.
— Guang MiIntelligence improvement is the only main line; intelligence itself is the biggest application
Guangmi compares AGI exploration to climbing a scientific mountain, and ultimately who can reach Everest. He repeatedly says one sentence: intelligence improvement is the only main line, intelligence itself is the biggest application. ChatGPT reached 3.5 with general generalization, unlocking conversational ability; Claude reached 3.5 Sonnet unlocking Coding ability, spawning Cursor; today everyone is unlocking Agent, Agentic. The climbing roadmap he draws: ChatGPT is just the first stop at the foot of the mountain, followed by Coding, Coding Agent, General Agent, AI for Science, Robotics. He believes AI for Science—scientific exploration—is that Everest: truly conquering cancer, or curing almost all human diseases.
— Guang MiRobot data collection is too inefficient, so it ranks after Science
Guangmi's attitude toward Robotics has changed. From first principles, today's Robotics Foundation Models or Research Labs are not essential enough in their approach. On data: GPT language models have a Scaling Law because there is the Common Crawl dataset continuously scraping internet data; but robot data collection is too inefficient—a person operating dozens of devices costs maybe tens to hundreds of yuan per hour, and to collect 100 million hours of effective data might cost hundreds of millions of dollars, making the cost of verifying Scaling Law very high. On algorithm structure: today there is no consensus on which underlying algorithm to use, and no architecture with general generalization has been found. He believes that in the future, as language models' multimodal capabilities become stronger, moving from the 2D world to the 3D world will be a relatively natural process.
— Guang MiOnline Learning may be the next paradigm-level route
Guangmi believes if there is another paradigm-level route in the future, Online Learning may be one. The core is not online real time, but letting the model autonomously explore and learn online, much like human survival—on the basis of survival and incentives, with ample curiosity, exploring everything, then abstracting and automating good workflows to form its own workflow. Under the current paradigm, the imaginable Online Learning is letting the model update smaller parameters in real time through interaction with users. But when to update memory, what to update, there is no reward today, and what goal the model should achieve is not well defined. He says Ilya's research on memory and multi-agent is relatively good.
— Guang MiThe profit pool division is unreasonable: Nvidia takes over 80%
Guangmi believes the profit pool division in today's value chain is very unreasonable. From NVIDIA to AWS to Anthropic to Cursor, NVIDIA takes almost over 80% of the profits, AWS takes 30%, Anthropic is losing money, and Cursor has negative gross margin. He judges that in the long term it will shift backward: AWS's profits will rise, then model and application profits will also rise. His confidence in the long-term value of model companies is growing stronger. He also mentions the economics of model training: OpenAI's tens of billions of dollars of investment provides a huge technological lever for all humanity, with a few thousand people leveraging the productivity of hundreds of millions, providing a huge deflationary force for today's inflationary world.
— Guang MiIn their own words · checked verbatim
I think coding might be a hand of the model. You see, Manus also built a virtual computer environment for the agent, a tool for the agent to operate the computer. I think within two years, agents will be able to operate digital environments like computers and phones, and their capabilities will absolutely surpass 99% of people.
我觉得coding可能就是模型的一个手吧 你看Manus也给agent搭了个虚拟的电脑环境 agent来操作电脑的工具 就是我觉得两年以内 agent能操作电脑和手机这种数字环境 它的能力是绝对超过99%的人的
Guang Mi11:15
What I've been repeating in my head recently is one sentence: intelligence improvement is the only main line, intelligence itself is the biggest application. So we still need to invest and think around intelligence itself.
我最近脑子里反复讲的就是一句话 就是叫智能提升是唯一的主线 智能本身就是最大的应用 所以我们还是要围绕智能本身去投入和思考
Guang Mi30:31
I think the competitiveness of organization and culture is the core competitiveness second only to computing power.
我觉得组织和文化的竞争力 是仅次于算力的核心竞争力
Guang Mi1:42:38
I think tech investment is not about mingling—you can't mingle your way to success. Many VC investors actually like to mingle in circles. I think mingling in circles is meaningless, the ceiling is very low. I think you still have to rely on creation.
我觉得科技投资不是混 能混出来的 很多VC Investor 其实喜欢混圈子 我觉得其实混圈子 没有意义 我觉得天花板是很低的 我觉得还是得靠创造
Guang Mi1:55:52
Figures
| Cursor's AR command count | Over 150, possibly 4 to 500 by year-end | 16:19 |
| Manus average tokens consumed per task | Seven to eight hundred thousand | 46:40 |
| Anthropic's latest valuation | 61.5 billion USD | 1:39:34 |
| Mira's new company valuation | 10 billion | 1:39:34 |
| Elia's company valuation | 31 billion | 1:39:34 |
| OpenAI's funding amount | 40 billion USD | 1:23:15 |
Glossary
- Pretraining
- The training stage that determines the model's intrinsic ceiling and can produce emergent new capabilities.
- Post Training
- The training stage after pretraining that strengthens and elicits the model.
- RL
- Reinforcement learning strengthens model capabilities through reward signals, but does not produce new capabilities.
- Tool Use
- The model's ability to call external tools to complete tasks, one of the key capabilities of an Agent.
- Long Context
- The model can process very long sequences of information at once, key for multi-step Agent tasks.
- Online Learning
- Letting the model autonomously explore and learn online, possibly the next paradigm-level route.
How to listen
Founders, investors and engineers watching the global large-model landscape; anyone who wants to understand the pretraining vs. Agent route split and how to judge competition among model companies.
The discussion of the China-US landscape and globalization after 1:54:32, which has lower information density.