The world is too loud. Read what matters.

张小珺·商业访谈录

AGI Is a Marathon: Physical Hardware Has Become the Biggest Bottleneck

Last year we thought adding GPUs and data would get us to AGI; this year we find that a single data center hits a wall at 32,000 cards, and US energy infrastructure was planned forty or fifty years ago—the AGI timeline has been stretched out by the physical world.

AGIComputeInfrastructureScaling LawRobotics
A frontline investor strings compute, energy, models, products and alliances into one AGI infrastructure map, dense with numbers, suited to anyone who wants to quickly calibrate their judgment.

The argument · tap a timestamp to hear it

6:32

AGI has gone from a sprint to a marathon

Early last year everyone thought AGI was a hundred-meter sprint and no one was prepared; this year it has become a marathon, and everyone has ample time to prepare. The change comes from a concrete discovery: last year we thought we could add GPUs and data indefinitely to reach AGI, but suddenly we found that GPU data centers and physical hardware are a huge bottleneck—a single data center now scales to 32,000 cards, and going beyond that requires breaking through many limits; US energy infrastructure was planned forty or fifty years ago, and the energy mix is very different, so suddenly adding a lot of electricity demand really can't keep up. So the biggest feeling this year is that physical hardware has become the biggest factor hindering the AGI timeline.

— Li Guangmi
23:41

The economy suddenly got two new taxes

One judgment is: model companies may be where value accumulates most thickly, just as mobile internet value accumulated in device makers and advertising platforms. Thus two taxes appear—Jensen Huang's GPUs collect a compute tax, and models collect an intelligence tax, suddenly adding two taxes to the economy and society. But Li Guangmi also gives the counterargument: if large models are electricity, the light bulb isn't necessarily made by the power plant. He leans toward large model companies being a research lab for basic discoveries; some labs have commercial capability and will produce top applications, but that tests organizational capability.

— Li Guangmi
26:34

GPT-4's entry ticket is 300 million dollars

Break training cost into electricity and cards: assume GPT-3.5 trains for 15 days on 500 H100s, about 250,000 kWh, which is 0.05% of the Three Gorges or Shanghai's electricity consumption; GPT-4 trains for 100 days on 8,000 H100s, about 26 million kWh, which is 5% of the Three Gorges or Shanghai's daily electricity consumption, and 2% of Texas's; if GPT-5 trains for 100 days on 32,000 H100s, it would be about 20% of the Three Gorges or Shanghai's daily electricity consumption. In money: each H100 sells for $30,000, plus peripheral equipment, 8,000 cards is over $300 million—$300 million is the entry ticket for GPT-4; a 32,000-card cluster means $1.2 to $1.3 billion.

— Li Guangmi
28:28

The bottleneck for a 10,000-card cluster isn't money

A 10,000-card cluster is standard, but money alone isn't enough. Every card must be connected, interconnection is very difficult, and the network topology is not one layer but three. Factors affecting difficulty include: finding land suitable for a GPU data center, stable and cheap electricity, interconnection and communication of data centers, and the reliability of cooling and operations. Resources are increasingly concentrated and converging; few customers can build large clusters, and it will converge to only four or five major customers. Things in the physical world are slower to transform than the digital world.

— Li Guangmi
40:38

Scaling Law hasn't slowed; compute just hasn't been pushed enough

GPT-4 publicly has 1.8T parameters, MOE architecture, about 13T data, trained for 100 days on 25,000 A100s. Assume the next generation has triple the parameters and triple the data, that's 9x compute; Jensen Huang's announced 32,000 H100 cluster plus optimization efficiency gains just matches; wanting 10x parameters and 10x data is 100x compute, clearly not enough. The conclusion is that Scaling Law has not slowed; if it seems slower, it's that compute and data haven't been pushed enough—GPT-3.5 to GPT-4 used roughly twenty to thirty times more compute, and from GPT-4 to the next generation we haven't yet pushed twenty to thirty times effective compute.

— Li Guangmi
45:43

This year you must surpass GPT-4, or you're out

2024 is the year of convergence for large model companies. The technical life-or-death line is that within this year you must surpass GPT-4, which requires a very strong team behind it; second- and third-tier model companies, including domestic ones, must surpass the best open-source models, otherwise their commercial value is relatively small. On compute, this year you must use a 10,000-card cluster, and few companies can do a good job with a 10,000-card cluster. The next 12 months will show whether a 100,000 H100 cluster is possible, with roughly $3-5 billion invested. Globally, the final survivors: in the US probably OpenAI, Anthropic, Google, and Musk's xAI; in Europe Mistral is good but it's uncertain whether Europe is an independent market.

— Li Guangmi
57:40

OpenAI doing products is forced by circumstances

There are three new understandings of OpenAI: the AGI timeline may be stretched; an AGI company shouldn't be too aggressive with products at first, but OpenAI is now very aggressive; and one begins to understand gradual unlocking. Why be aggressive with applications? Li Guangmi guesses: if AGI takes 10 years, each year requires several billion or even 10 billion dollars of investment, so commercialization is needed, and sustained healthy cash flow is needed to support AGI; relying purely on financing it's hard to raise that much money, and you can't just depend on Microsoft. He also judges that OpenAI is harder on the 2B enterprise side, because enterprise customers care about trust, and Microsoft is too deeply trusted by enterprise customers; a large part of OpenAI's 2B value may be taken by Microsoft.

— Li Guangmi
1:07:20

Silicon Valley VC's big three: coding, agent, robotics

Silicon Valley VC's investment theme this year is the big three: coding, agent and robotics. But Li Guangmi is skeptical of all three: coding is certainly within the core range of large model companies and Microsoft, the core capability comes from model companies, these coding startups won't train their own large models, and it's uncertain how much value the optimization layer on top has; model companies may be very aggressive with agents, because the added value is high, it's model-level capability, model-level application, model-level agent. He leans toward short-term value still accumulating in the models themselves. Robotics is the first choice for many researcher startups, because it's easier to tell a story.

— Li Guangmi

In their own words · checked verbatim

The biggest feeling this year is that physical hardware has become the biggest factor hindering the AGI timeline.

今年最大的一个感受就是 物理硬件成为阻碍AGI的一个时间表的最大因素了

Li Guangmi6:45

One is Jensen Huang's GPU collecting tax, right, one is the model collecting an intelligence tax, suddenly adding two more taxes to the economy and society.

就一个是老黄的GPU收税 对吧 一个是模型收一个智能税 突然给经济社会又加了两道税吧

Li Guangmi23:41

Its average is not about money. Right, a 10,000-card cluster is standard, standard, having money is not enough, it's very hard.

它的平均不在钱上 对 万卡集群这个是个 嗯 标配 标配 它是有钱是不够的 很难的

Li Guangmi28:28

People's expectations can fly very fast, but the physical world, I think, can't keep up.

就是人的预期可以飞得很快 但是物理世界我觉得是跟不上的

Li Guangmi32:34

If I must give a conclusion, I think Skill up has not slowed down. If it has slowed, I think it's that compute and data haven't been matched enough.

如果非要说一个结论 我觉得Skill up是没有减速的 如果说变慢了 我觉得就是算力和数据没对够

Li Guangmi41:40

The carbon-based body still has many limitations. For example, compared to a large star, a person's throughput is limited, memory is relatively short, and one can't work long hours, energy is also a problem.

探机肉身 还是有很多局限的 你比如说 相比大猛星来讲 人的吞吐量是有限的 记忆也是比较短的 也没办法长时间工作 精力也有问题

Li Guangmi1:15:47

Humanity has two more tax-collecting cornerstone companies: one is chip compute water, one is model intelligence water.

人类又多了两个 收税的人类基石公司 一个是芯片算力水 一个是模型智能水

Li Guangmi1:20:09

Figures

Cards per single data center32,000 cards6:32
GPT-3.5 training electricityAbout 250,000 kWh (500 H100s for 15 days)26:34
H100 price$30,000 each27:28
GPT-4 entry ticket$300 million27:28
Microsoft's investment in OpenAI$13 billion30:33
Nvidia GPU annual shipmentsPossibly 4 million this year30:33

Glossary

Scaling Law
The rule that model capability improves as parameters, data and compute scale up.
Agent
A model application form that can autonomously call tools and complete multi-step tasks.
Sovereign AI
Governments using defense and other budgets to procure GPUs and build domestic AI capabilities.
T5T
Nvidia's internal habit of every large group sending out the five most important things every two weeks.
Killer App
An application that brings large-scale users and revenue and defines the platform's value.

How to listen

Who it's for

Founders, investors and engineers following AGI progress, who want to use one episode to calibrate their judgment on compute, energy, models and products.

Skip

After 1:15:47, the chat about carbon-based and silicon-based life can be skipped.