The world is too loud. Read what matters.

张小珺·商业访谈录

Only three companies have replicated GPT-4, and the window closes in 2024

Globally only OpenAI, Anthropic and Google have produced GPT-4-level models, and how 2024 plays out basically decides the long-term landscape — if you can't catch up in the next 12 months, it's very hard to flip the board afterwards.

Large ModelsSilicon ValleyFinancingScaling LawAGI
An investor who spent the whole year physically in Silicon Valley and backed two overseas large-model companies lays out financing, talent, cost, camps and endgame predictions as one continuous line. Information density is high, but the second half is worth more than the first.

The argument · tap a timestamp to hear it

4:04

There is only one core secret: data and tokenizer

Asked what the single most core secret is at OpenAI, Anthropic and Google, the answer is that in the short term it's data — the mix of pre-training data, how the tokenizer is built; and if you could keep only one capability, it would be reasoning. This judgment pulls the ‘model race’ back from a compute narrative to data engineering: globally only two or three hundred people actually know GPT-4's data secret, almost all of them at the top few model companies, and any other company trying to figure it out needs at least several hundred, even several thousand, sufficient experiments — a few tens of thousands of GPUs is a necessary condition.

— Li Guangmi
14:10

The final circle for replicating GPT-4 has only three players

Taking replication of GPT-4 as the threshold for entering the final circle, globally only OpenAI, Anthropic and Google have produced GPT-4-level models. On the timeline: OpenAI did GPT-4 a year ago, Anthropic half a year ago, Google can only deliver next month, and other teams worldwide may need another 6 to 12 months. Google threw the whole company at it for a year and only barely got close. Three high-potential dark horses were named: Musk's XAI, Character built by Norm, a core contributor to Transformer, and Bidance.

— Li Guangmi
16:10

Only two or three hundred geniuses can build large models

The talent barrier in large models is extremely high; globally the genius researchers who can actually make a real, major contribution to large models may number only two or three hundred, of whom over a hundred are at OpenAI, twenty or thirty at Google, and Meta, AWS and NVIDIA may have none. And the clustering effect is strong — these kinds of people and this kind of culture matter enormously, which is why there's less optimism that other giants can do large models well on their own. The test isn't the CEO, it's whether the model company has at least one genius scientist: OpenAI has Ilya, Anthropic has Dario, and the CTOs of Runway, Ideogram and Pika Labs are all key figures in their respective directions.

— Li Guangmi
21:11

Next-generation training cost jumps from two or three hundred million to a billion

If the training cost of Claude 3 and GPT-4.5 is 200 to 300 million dollars, then the next generation in 2025 and 2026 could cost at least 1 billion dollars, even 3 to 5 billion. Training cost splits into two parts: three quarters goes to experiments (repeated trial and error on small-scale models), one quarter to the final big training run — like a rocket launch. The publicly rumored figure for GPT-4 was 22,000 A100s training for 100 days, with the pure big-training cost at roughly 80 million dollars, but the biggest cost is the earlier experiments. 70 billion parameters is a dividing line: below it you can tolerate many errors, above it every step up makes training difficulty rise exponentially.

— Li Guangmi
30:23

Once 2024 plays out, the landscape is set

How 2024 plays out basically decides the rough landscape; the window may be the next 12 months, and if you can't catch up in the next 12 months, it's very hard to flip the board afterwards. The most idealized landscape is that probably only one company remains — the most advanced model is also the cheapest, so there's no reason to use a second one; but because of camp rivalry, in practice it may be two or three. The camps: Microsoft and OpenAI on one side, Amazon and Google together backing Anthropic on another, Google as its own camp, and Apple and Tesla representing the device camp. Meta isn't necessarily a large-model company; it's a company using AI to do its own business well.

— Li Guangmi
37:35

Cost is the overlooked invisible competitive edge

Large models going forward have two main threads: the visible line is rising intelligence capability, and each increment unlocks some new applications; the invisible line is falling cost — model training cost fell 4 to 5 times over the past 18 months, inference cost fell by a factor of 10 over the past 18 months, and another two or three rounds of optimization shouldn't be a problem. These two threads determine the scale of the AI-native application explosion. The core of cost reduction is an engineering problem — GPU utilization, architecture optimization, precision tuning. If you can make cost extremely low while the model is still competitive, that's an extremely strong core competitive edge, very much like chips, and a certain scale effect is already visible in the leading companies.

— Li Guangmi
46:37

Over 80% of model companies will be acquired

Silicon Valley model companies today are more like a research lab; apart from ChatGPT's unexpected breakout, the business model is still unclear. Even for Silicon Valley large-model companies, an independent IPO may be very hard — 80% to 90% will most likely be acquired, so large-model companies still need to hug a big leg. Silicon Valley VCs have almost all missed large-model investment, just as they missed Tesla and SpaceX — this is a problem of two products not matching: a company like a large-model one, high-risk, high-investment, with an unclear business model, doesn't fit the typical investment profile of the VC financial product.

— Li Guangmi
50:41

ChatGPT shows no network effect or data flywheel

The internet is about network effects, data flywheels and scale effects, but large models and AI today don't seem to show these yet. ChatGPT may be more like a consumer product; it only knows the distribution of some user queries, which can better guide training and which data matters or doesn't, and it can also distill smaller models to serve head queries, but it's not yet a product with a strong data flywheel or network effect. The core of the mobile internet era was that the world gained four or five billion more users and phones could collect more data for machine learning and recommendation; in this wave, one of the most invisible core competitive edges may be cost.

— Li Guangmi
54:53

At the start of the year I underestimated GPT-4's difficulty and overestimated the speed of applications

Looking back on the year, at the start I underestimated the difficulty of achieving GPT-4 and overestimated the speed of the application explosion. Today there are still too few truly AI-native products, just the top few — ChatGPT, Character, Perplexity — and we need to wait another one or two generations of models before more native products appear. Enterprise exploration of large-model use cases still has few successes, with only Microsoft and Adobe relatively aggressive. Large models are still early today, very much like chips: you have to wait for chip capability and cost to iterate another two or three generations before consumer electronics on top slowly explode.

— Li Guangmi
1:11:59

The incremental GDP created by AI is 5 to 10 times that of the internet

A projection for the next 20 years: the direct incremental GDP created by AI may be 5 to 10 times larger than the incremental GDP the internet created over the past 20 years. The math: if this wave of AI replaces 1 billion white-collar workers, each earning 30,000 to 50,000 dollars a year, that's a 30 to 50 trillion dollar market size; if global GDP doubles, from today's 96 trillion dollars to 200 trillion, an increase of 100 trillion, and AI takes 10% to 20%, that's 10 to 20 trillion dollars in revenue, multiplied by a multiple of 10, and many big companies will be born. Also, data center electricity today accounts for 2% to 3% of humanity's total energy, and rising to 10% to 20% in the future is quite foreseeable.

— Li Guangmi

In their own words · checked verbatim

So I think large models today are a hundred-billion-dollar BET for humanity.

所以我觉得大模型今天是人类 一个千亿美金的BET

Li Guangmi45:35

Everyone treats ChatGPT and Character as applications. I think that's noise — actually the two of them are model companies.

大家都把chatGPT和character当应用 我觉得这就是噪音 其实它两个是模型公司

Li Guangmi1:26:05

Jobs and Musk seemed to have no friends in Silicon Valley, but Sam is friends with everyone in Silicon Valley.

乔布斯和马斯克 好像在硅谷没有朋友 但Sam在硅谷所有人都是朋友

Li Guangmi1:28:07

When you see companies that impressive in Silicon Valley, scientists that brilliant, all charging forward with such belief, I think you see light and hope in their eyes.

当你看到硅谷那么牛逼的公司 那么天才的科学家都一往无前很相信 我觉得从他们的眼睛里是看到了光和希望吧

Li Guangmi1:30:07

Figures

OpenAI annualized revenueOver a billion dollars, possibly five or six billion next year12:09
ChatGPT stable MAUOver 200 million4:04
ChatGPT share of Chat traffic70% to 80%12:09
GPT-4 training scale22,000 A100s training for 100 days40:33
Model training cost declineFell 4 to 5 times over the past 18 months38:31
Model inference cost declineFell by a factor of 10 over the past 18 months38:31
Global large-model investmentAt least 100 billion dollars to be spent over the next three to five years44:35

Glossary

Scaling Law
The empirical rule that model capability improves as parameters, data and compute scale up; there is still no theoretical support for it.
Tokenizer
The method for cutting text into units the model can process, directly affecting training results.
MoE
A model architecture, proposed by Norm, that lets different expert networks handle different inputs.
Diffusion Model
The mainstream approach for image generation; adding a time dimension makes video generation possible.
AI native
Products designed from model capability outward, rather than layering AI features onto an old product.

How to listen

Who it's for

Founders, investors and engineers tracking the large-model race, especially anyone trying to understand the 2024 landscape window, the endgame for model companies, and the cost curve.

Skip

1:15:00 to 1:17:00 covers car companies and autonomous driving, weakly connected to the main thread.