Scaling laws were discovered first at Baidu, but GPT was built by an American company
MiniMax founder Yan Junwei (严俊杰) argues that the scaling law underpinning large language models was discovered at Baidu in 2014—yet it was a history-free American startup, not a Chinese company, that actually built GPT.
The video won't play here. Listen to the audio instead:
The argument · tap a timestamp to hear it
China's AI spending is a fraction of America's but only trails by 5% technologically
Yan Junwei says that of the four American companies truly capable of building large models—Google, OpenAI, Anthropic, and xAI—their valuations are roughly 100 times those of comparable Chinese startups, and revenue is similar. Yet their R&D spending is 50–100 times higher, while technological advantage is only about 5%—and that gap has been shrinking over the past year. He attributes this to China's talent density and engineering innovation forced by compute constraints, not just copying American methods.
— Yan JunjieScaling law abandoned a decade of accumulated academic theory
When MiniMax was founded, its first insight was that building large models requires abandoning past experience. Previous AI valued complex mathematics and elegant theory, but this generation is defined by scaling—using the simplest possible method to let performance improve continuously with data and compute. Yan admits much of his prior decade in AI research "turned out to be useless"; what truly matters is finding a way to scale everything: team organization, technical decisions, and business.
— Yan JunjieThe model itself is the product; apps and interfaces are distribution
When Luo Yonghao asked whether product managers would become obsolete, Yan's answer reframed the entire category. In the large-model era, the actual product is the model itself, because what users pay for is intelligence—intelligence that the model computes. Traditional product artifacts—apps, interfaces—function more like distribution channels. MiniMax's own app is essentially one in-house store; partner integrations are another channel. This belief guided the company to place its algorithms team at the core.
— Yan JunjieBaidu discovered the core principle behind large models—six years before OpenAI
During his internship at Baidu in 2014, Yan had access to what was then China's only large-scale GPU cluster, and as an intern he could command a third of it for experiments. Only later did he learn that Dario Amodei, now CEO of Anthropic, was also at Baidu working on speech, and that it was there he discovered the scaling law—the central principle of large models—six years before it captured industry attention. That GPT was built in America, not China, he sees as partly chance and partly China's loss.
— Yan JunjieAI-first strategy and internet playbooks cannot survive side by side
Early on, MiniMax considered the obvious approach: since the company needed to serve both 2C users and international markets, why not apply China's most proven international 2C playbook—mobile internet tactics—adding AI talent plus internet expertise. Experimentation proved this wrong. Building a great product through model capability and building one through replicated internet classics may each work, but they cannot coexist. The driving principle must choose one path. This realization tormented the team for months before they committed to technology-first, despite its higher cost and risk.
— Yan JunjieDeepSeek forced the company to discard OKRs and KPIs entirely
DeepSeek's explosion around Chinese New Year hit MiniMax hardest not as competitive pressure but as a signal that organizational structure must transform. The algorithms and infrastructure teams cannot be separate; all people must serve a single optimization target, not individual metrics for each person or team. When they attempted OKRs, they found they don't work at all; KPIs work even less, because whether the next model generation succeeds is fundamentally uncertain. The result: the company now lacks management tools at all, and some who disagreed with this approach left.
— Yan JunjieHalf the company once argued they should abandon language models
For roughly two years, MiniMax's language model generated zero commercial value for the business while consuming the most compute and talent—one of the company's most painful conflicts. Yan says at least half the staff argued they should drop the language-model line entirely. A similar battle over abandoning domestic Chinese operations played out alongside it: international markets showed far better ROI, but Yan insisted China is ultimately the largest single market and must stay the course until the business model proves itself.
— Yan JunjieBest-in-class models capture premium use cases, amplifying revenue gaps 100-fold
Yan explains how the revenue gap between Chinese and American large-model companies grows even wider. Chinese models trail by only 5%, but in the most valuable, high-end scenarios that demand top-tier capability, customers will use only the best model. This effect magnifies the addressable market by 10x and prices by another 10x—multiplying to a 100x business gap. This is precisely where he believes Chinese companies must focus in the coming year: gaining more than single-digit share of the global mainstream language-model market.
— Yan JunjieIn their own words · checked verbatim
For instance, top American companies have valuations roughly 100 times those of comparable Chinese startups, revenue is also about 100 times, but their technology only leads by around 5%
比如说美国最好的公司的估值 是中国 可能创业公司里面 大概是估值是100倍 啊 然后收入呢 基本上也是100倍 但是技术呢 可能就领先5%
Yan Junjie4:01
What really matters is finding a way to organize our team that can scale
真正重要的东西 就是找到一种能够 scaling规模化的方式 来组织我们的团队
Yan Junjie11:05
For anything that can be quantified, we believe a model will necessarily outperform humans or reach the level of the best human
对只要是一个东西能被量化 我们认为就是模型它 一定会强于人 或者一定是能到最好的人的那样的水平
Yan Junjie30:14
You could improve the product or business either through model capability or by copying mobile-internet classics—both might work, but they can't coexist
比如说通过模型能力来 让这个产品变好 或者业务变好 和是说 比如说通过把这种 移动互联网的这些经典的东西 给复制过来 让它变好 这两个东西 有可能都是对的 但这两个东西 其实是没法共存的
Yan Junjie1:21:32
We tried OKRs and found they don't work at all; KPIs work even less, so we need entirely new tools—which means we have no tools
我们试图用OKR 我发现根本行不通 那KPI更行不通 更行不通 所以我们就 需要全新工具 所以我们没有工具
Yan Junjie2:45:29
This was an extremely painful situation; at least half the company believed we should not pursue language models
这是一个非常痛苦的一件事 可能会有 公司里面至少有一半的人认为 我们应该不做语言模型
Yan Junjie3:06:37
Figures
| US top-tier LLM company valuations vs. Chinese startups | ~100x | 4:01 |
| China-US cutting-edge LLM technology gap | ~5% | 4:01 |
| China-US LLM R&D spending gap | ~50–100x | 4:01 |
| Global language-model market size | ~$20 billion | 12:06 |
| Global image-and-video model market size | ~$2–3 billion | 12:06 |
| Global audio-model market size | ~$300 million | 13:06 |
| Global music-model market size | tens of millions of dollars | 13:06 |
| GPT-3 R&D cost | ~$100 million | 1:00:24 |
| MiniMax team size | 400+ | 1:44:48 |
Glossary
- Scaling Law
- The principle that model capability improves continuously as data and compute increase; the central assumption behind large-model development.
- AGI
- Artificial general intelligence—a system capable of performing any intellectual task a human can.
- Exploration/Exploitation
- In reinforcement learning: exploration (trying diverse strategies) versus exploitation (optimizing known-good strategies); the trade-off that shapes learning.
- MOE
- Mixture of experts—a model architecture that routes different tasks to specialized sub-networks.
- OpenRouter
- A developer platform that provides a unified interface to call large-model APIs from multiple providers.
How to listen
Founders and investors seeking to understand Chinese AI companies' real cost structures, technology trade-offs, and organizational logic.
32:00–52:00 covers childhood and education—heavily personal in tone, with low density of business and technical information.