The first to build AGI will inevitably be a monopolist; China's only path is the ecosystem
The pursuit of an AGI that crushes all rivals first is, at bottom, an ambition for commercial hegemony — whoever holds that brain will invent the means to preserve the monopoly. Chinese companies can't win this money-burning game, so they have to drive inference costs down first and use the ecosystem as their moat.
The video won't play here. Listen to the audio instead:
The argument · tap a timestamp to hear it
If inference costs don't fall, the application ecosystem can't take off
01.AI's most important technical decision in 2024 was to commit unreservedly to the mixture-of-experts (MoE) route and build models that are highly capable, fast and cheap. Li Kaifu's insight: the whole industry's detonation will happen on the application side, and the application side needs very cheap inference to detonate. He ran the numbers — applications will get harder one after another; a chatbot can tolerate spitting out one character at a time, a search engine has to be faster, and something replacing Douyin has to be scrolled for three to five hours. Usage intensity and model difficulty keep escalating, and multiplied together, inference costs falling tenfold a year still wouldn't be enough.
— Kai-Fu LeeVertical integration is the only way to drive costs down
Li Kaifu doesn't buy the idea of lying flat and waiting for GPU and memory prices to fall — enjoying a tenfold price cut a year and jumping out to do applications at the right moment. He cites the iPhone: Apple didn't wait for capacitive screens, multi-touch, the soft keyboard and the music API to become industry standards; it folded them together one by one, because waiting might have taken five to seven years. So 01.AI chose vertical integration: hardware to software, model to application, all combined and optimised together. Concretely, that means converting compute into memory — designing three tiers of memory around the GPU: HBM, the CPU's RAM, and SSD. Spending 10% more on memory might cut GPU compute by 40% to 50%.
— Kai-Fu LeeBuilding the world's best model and building applications don't connect
Li Kaifu uses Character.AI as an example: founder Noam Shazeer was single-mindedly pursuing AGI and first built the character.ai product, but the two things conflict — an AGI model is a sledgehammer to crack a nut and doesn't suit characters. The company struggled to find its place and was eventually acquired by Google, with the AGI people going back to Google and the app people staying at Character. He explains the mechanism: people chasing the world's number-one AGI have a mix of idealism and arrogance, believing that once AGI is achieved it crushes all competitors, that there may be no ecosystem above it, that every app is just a thin shell — veneer in English — and all the value sits in the AGI. Li Kaifu says he isn't that kind of person.
— Kai-Fu LeeThe first to build AGI will inevitably be a commercial-hegemony monopolist
Li Kaifu reasons it through: suppose AGI has thinking and reasoning, invention and creation, independent thought and self-learning far beyond humans — then it can help you invent new things: invent a business model that knocks down the other large-model companies, design a PR strategy that makes everyone trust you, even design a cyberattack that paralyses your competitors. So whoever builds AGI first and crushes their rivals is of course pursuing a technological ideal, but is also a commercial-hegemony monopolist, and it will become the ambition of an ultimate monopolist. This is far bigger than the old talk of Microsoft's monopoly or Google's monopoly, because in the past it was hard for Microsoft to make Windows wipe out the Mac, but this is a brain.
— Kai-Fu LeeToday's AI ecosystem is an inverted triangle
Li Kaifu draws an ecosystem diagram: a healthy ecosystem should have chips earning the least, platforms earning quite a lot, and applications earning the most — all applications together earning more than the platform, which is how PC, internet, mobile internet and cloud all worked. But today's AI ecosystem has chips/GPUs at $75 billion, cloud vendors at $10 billion, and application vendors (the ChatGPTs of the world) at only $5 billion — a complete inverted triangle. If the inverted triangle persists, applications won't spring up like bamboo shoots after rain, users won't get the benefits, and app builders will struggle to reach PMF quickly, make money and raise funding; the positive cycle can't get turning.
— Kai-Fu LeeOpenAI is still holding many cards; don't underestimate it
Li Kaifu has just come back from Silicon Valley, and a consensus among many people he met is that OpenAI is still holding a lot of good things it hasn't released. GPT5 training hasn't gone smoothly, but to raise money it tossed out an O1; it still has many cards in hand and isn't in a hurry to play them — because every time it plays one, global tech companies including China's look at what it played and guess and copy, and in the end even if they can't match it they get to 80 or 90 percent. So it wants to hold back until AGI is nearly in sight. Li Kaifu thinks O1 itself didn't bring that big an improvement in reasoning and understanding; the most impressive thing is the inference-time scaling law — the straight line it drew opened up another gold mine.
— Kai-Fu LeeNvidia's advantage may wobble as training shifts to inference
Li Kaifu goes company by company on the overseas giants: Nvidia is the biggest beneficiary now, but the problem it may face later is whether its advantage can persist when most GPUs are no longer doing training but inference. Meta is the biggest disruptor — whatever it can't win at, it open-sources. This time it looks fairly successful: make some money from ads, open-source to stake out a position. Its technology lags, but open-sourcing when you can't beat someone has already worked twice (the first being TensorFlow versus PyTorch). Microsoft is best positioned, able to attack and defend, but its challenge is that its own models have never been good; long term, if it can't build models and falls out with OpenAI, that's a challenge.
— Kai-Fu LeeGoogle's search faces enemies on two fronts
Li Kaifu says Google is rather sad: in theory it should be the strongest — the best large-model papers were done by Google, the best reinforcement learning was done by DeepMind — but put together they don't seem to have produced much lethality. Search is challenged on all sides: large models mean users ask ChatGPT first when they have a question, taking away some volume; more seriously, many users who want to buy something go straight to Amazon, so commercial search isn't happening on Google anymore. Add the dilemma: do you put large models into search? Replacing search would dismantle the ad business; making two entry points is burying your head in the sand; having both coexist makes it a neither-fish-nor-fowl — give an overview but no answer, plus a pile of links and ads.
— Kai-Fu LeeIn their own words · checked verbatim
Because if you're doing AGI, you have a kind of idealism and arrogance coexisting — that once I've built AGI I crush all my competitors, that there may be no ecosystem above me, that every app I have may be just a thin shell, what we call in English a veneer.
因为你要做AGI的话 你就是有一种理想和傲慢 两者共存 就是我做成了AGI 我就碾压所有的竞争对手了 我上面就未必有什么生态系统了 我可能每一个APP就是薄薄的一层壳 我们英文叫veneer 贴皮
Kai-Fu Lee28:28
So the first to build an AGI that crushes its rivals — of course it's a technological ideal, but it's also a commercial-hegemony monopolist, and it will become the ambition of an ultimate monopolist.
所以第一个做出AGI碾压对手的 当然是一个技术的理想 但是他也是一个商业霸权垄断者 而且会成为一个终极垄断者的一个野心
Kai-Fu Lee41:25
We'd see a healthy ecosystem where chips earn the least money, platforms earn quite a lot, and applications earn the most — the platform itself may earn more than any single application, but all the applications together earn more than the platform.
我们会看到一个良性的生态 应该是芯片赚最少的钱 平台赚蛮多的钱 应用赚最多的钱 平台本身可能比任何一个应用都赚钱 但是所有的应用加起来是比平台赚更多的钱
Kai-Fu Lee45:44
What does "usable by everyone" mean? If you build OpenAI in America and don't let Chinese people use it, then it's not usable by everyone, so you don't dare write it into your vision.
人人可用什么意思呢 你在美国做了OpenAI不给中国人用 那就不是人人可用 所以你不敢把它写到你的vision
Kai-Fu Lee1:12:11
I think OpenAI is still holding a lot of good things it hasn't released. We must not underestimate it. Its GPT5 training hasn't gone very smoothly, but to raise money it tossed out an O1. It still has many cards in hand, and it's not in a hurry to play them.
我觉得OpenAI还藏了很多好东西 没有放出来 我们千万不要低估它 他GPT5训练的不是很顺利 但是他为了融资就丢了一个O1出来 他手中还有很多牌的 他不急着出这些牌
Kai-Fu Lee1:14:14
Actually that's absolutely not the case. We measured Perplexity's hallucinations and they're actually quite high, but users just feel that with a citation, once they see it they're reassured — they don't actually click through and they trust it. This is a very interesting and very worth-learning user-experience trick.
其实绝对不是的 我们衡量过Perplexity的幻觉 其实还挺高的 但是用户就觉得是Citation 看到了就放心 你不去真的点就信任他了 这个是很有意思 很值得学习的一个 用户体验的Trick
Kai-Fu Lee1:33:25
Figures
| 01.AI's product revenue this year | possibly five or six million dollars | 33:31 |
| GPT-4O vs 01.AI model inference cost | 5 months ago $10 per million tokens, today down to one mao four per million tokens | 48:47 |
| Expected time for AGI | 7 years from now | 49:43 |
| Today's AI ecosystem chip/GPU scale | $75 billion | 45:44 |
| Today's AI ecosystem cloud vendor scale | $10 billion | 45:44 |
| Today's AI ecosystem application vendor scale | $5 billion | 45:44 |
| LAMA405B inference cost vs 01.AI | about 20x | 1:20:18 |
| Revenue Google generates per search | 1.6 cents | 1:32:23 |
Glossary
- veneer
- Li Kaifu's term for how every app above AGI is just a thin shell, with all the value in the underlying model.
- Service as Software
- A new Silicon Valley phrase for delivering digital employees through a software interface and charging by one person's workload.
- TCPMF
- Li Kaifu's two extra letters on top of PMF: you also have to judge how strong the technology needs to be, who can build it when, and when costs will be low enough.
- HBM
- The fastest tier of memory next to the GPU, and the place where the most frequently used data sits in Li Kaifu's three-tier memory architecture.
How to listen
Founders and investors watching the AGI geopolitical landscape, large-model companies' strategic choices and overseas giants' positioning — especially anyone trying to judge when the application layer will take off.
The opening 2:03 to 10:20, on Hinton's Nobel Prize and Li's student years, can be skipped; it's mostly reminiscence.