The world is too loud. Read what matters.

张小珺·商业访谈录

Yang Zhilin: Model Companies Are Refineries; the Value Is in Chemical and Car Companies

In business history, the first wave of breakout companies is almost never the winner of the next stage. Model companies are refineries that extract base oil; the real value is in chemical and car companies — which is why Moonshot AI does only to C, and only to C.

Large ModelsAGIStartup FinancingLong ContextOrganization Design
Yang Zhilin rarely lays out technical judgment, organizational design and the financing window as one line; the second half, on Sora, world models and why open source can't catch closed source, has the highest density of reasoning.

The argument · tap a timestamp to hear it

6:05

An AI Lab can produce a Transformer, but not GPT

Yang Zhilin breaks ‘breakthroughs happen in industry’ into two layers: first, the inevitable law of field development, as exploratory research gradually shifts to an industrialization process; second, the limits of organizational form. Google Brain is essentially a research organization inserted into a large company; this organizational form can explore new ideas, but it is very hard for it to produce a great system — ‘it might be able to produce a Transformer, but it might not be able to produce GPT.’ He judges that the scientific research system and the education system will later shift their functions, mainly to cultivating talent. This is also how he explains why AGI needs a new organization: technology determines the mode of production, the organization must adapt to the technology, and if it doesn't match, it cannot produce effectively.

— Yang Zhilin
13:12

Release yourself from infinite carving

This is the most important lesson he learned at Google. The picture he gives is: there are ten roads in front of you; most people consider how to brake when there is a pedestrian ahead on the road they take, but which of the ten roads to take is the most important thing. At the time, the problem in this field was that the focus was wrong — lowering loss further on a dataset of one or two million tokens, inventing all kinds of bizarre architectures and regularization methods, were all carving techniques; they improved on the dataset but did not see the essence of the problem. The essence is to find first principles: as long as a structure is general enough (all problems can be put into it for modeling) and scalable (put in enough compute and it gets better), you should not do excessive carving at the upper layer. The line he quotes is: if you can solve a problem with scale, don't solve it with a new algorithm.

— Yang Zhilin
22:17

The first financing window was only one month

Yang Zhilin describes the timing very concretely: they began concentrating on the first financing round in February, and this window was very short — delay to April and there was basically no chance, but doing it in December or January also had no chance, because at that time there was still covid and people had not yet reacted. In China it exploded in February; in November of the previous year there was still little reaction, so the real window was one month. In the US he did a precise calculation every night: how many flops that corresponded to, how much training cost, how much inference, how many users; the conclusion after calculating was that they had to get at least 100 million USD within a few months. At the time many people thought it might not be possible to raise that much; later it proved possible, and even more. After the financing window came the hiring window in March and April.

— Yang Zhilin
27:25

Hire geniuses first, then raise the floor

Yang Zhilin says hiring is the core of the core; if the people are wrong, even the most superb management skills are useless. There are very few people in the world with real AGI experience, so their early profile was very focused: directly find the matching genius — someone with the ability to operate on models and direct experience training at super-large scale can build it quickly. For a long time the team was thirty or forty people, now eighty, deliberately pursuing talent density and not wanting too many people. His reference is that Google now has several thousand people doing this; if it were cut to one hundred or fifty, it might succeed quickly. Only later did they begin to fill in product, operations, leader-type talent and animals who can push things to the extreme — filling in the floor.

— Yang Zhilin
32:27

Long context is the memory of the new computer

Why do long context first? Yang Zhilin's analogy is: it is the memory of the new computer; the memory of old computers grew by several orders of magnitude over the past few decades, and the same thing will happen to the new computer. It solves two essential problems — generality and personalization. On generality, when you have a sufficiently lossless long context, in a multimodal architecture you may not even need a tokenizer; you can put the raw signal directly in. On personalization, the core value of AI lands on personalized interaction, and personalization is not achieved through fine-tuning; it is a process defined by a very long context that cannot be replicated. He stresses that this technology has been worked on for more than half a year, not seeing the trend and then gathering two teams to develop quickly; the two are very different.

— Yang Zhilin
42:36

The next two milestones: a unified world model and evolution without human data

Yang Zhilin gives two of the most important technical milestones: first, a truly unified world model that can unify different modalities, a truly scalable and general architecture; second, being able to let AI continue to evolve without human data input. His judgment is that many of the reasoning and agent problems discussed today are products after these two problems are solved; there may still be some carving to do, but there will be no fundamental blocker. He also uses this standard to measure product value: if only ten or twenty percent of a product's core value comes from AI, then the thing does not hold; intelligence is always the core incremental value.

— Yang Zhilin
47:36

Open source can't catch closed source, because open source itself is still centralized

Yang Zhilin's reason is not how many resources there are, but that the development method has changed: previously everyone could contribute to open source, but now open source is essentially still a centralization that requires the concentration of resources, talent and capital, so closed source is better; it will be a process of consolidation. He adds an angle: if today there were a very leading model still open-sourced, it would most likely be unreasonable. This directly supports his rejection of ‘using open-source models to build a super app’ — without technical capability, finding a super app is impossible.

— Yang Zhilin
48:36

Rushing to find PMF will get you dimensionally reduced

In response to views that only invest in applications, such as ‘10 people can't find PMF, 100 people can't find it either, the key is whether you can find PMF,’ Yang Zhilin's response is that the ceiling of AGI is far higher than everything seen now, and one must think in terms of 10 to 20 years. The historical case he gives is companies that previously did dialogue systems, customer service systems, slot filling; some had decent scale, but all were dimensionally reduced, because no one will use that kind of technology again. And today is still in the process of change — context is getting longer and longer, instruction following is getting stronger and stronger, so the difficulty of customizing a customer service system will get lower and lower and accuracy higher and higher; therefore this is a 100% dimensional reduction. He acknowledges that focusing only on applications has commercially viable opportunities, but the biggest opportunity is not there.

— Yang Zhilin

In their own words · checked verbatim

It might be able to produce a Transformer, but it might not be able to produce GPT.

它可能能产生Transformer 但可能产生不了GPT

Yang Zhilin6:05

If you can solve a problem with scale, don't solve it with a new algorithm.

如果你能用scale解决的问题 你就不要用新的算法去解决它

Yang Zhilin14:13

Only if it is a disruptive thing does it deserve the three letters AGI; otherwise everything we say today is meaningless.

你就是一个颠覆性的东西 它才配得上AGI这三个字 否则我们今天所有说的这些事情都没有什么意义

Yang Zhilin43:36

If you rush to find a PMF, you will very likely find in the end that you have been dimensionally reduced again.

如果是着急的去找一个PMF 你很有可能最后发现 又被降维打击了

Yang Zhilin48:36

Figures

Moonshot AI team sizeEarly thirty or forty people, about 80 at the time of the interview28:26
Total of Moonshot AI's two financing roundsClose to 2 billion RMB26:25
Financing target Yang Zhilin calculated in the USAt least 100 million USD within a few months22:17
First financing windowAbout one month (February)22:17
Price fluctuation of a single machine260 one day, 340 the next, then fell back a couple of days later30:27
Yang Zhilin's paper citationsMore than 22,0001:02

Glossary

scaling law
As long as you put enough compute into the model, the results get better.
long context
The length of text a model can process at one time; Yang Zhilin calls it the memory of the new computer.
PMF
A product finding a scenario that is truly needed by the market.
DiT
The architecture used by Sora; Yang Zhilin believes it is not yet a general architecture.
next token prediction
The only basic principle of language models that has been verified to work.

How to listen

Who it's for

For founders and investors judging the large-model startup window, organizational form and the to-C route, especially those agonizing over whether to chase PMF.

Skip

After 1:15, the China-US ecosystem and 2024 predictions are fairly general; you can listen only to the Sora technical judgment part.