Yang Zhilin: Model Companies Are Refineries; the Value Is in Chemical and Car Companies
In business history, the first wave of breakout companies is almost never the winner of the next stage. Model companies are refineries that extract base oil; the real value is in chemical and car companies — which is why Moonshot AI does only to C, and only to C.
The video won't play here. Listen to the audio instead:
The argument · tap a timestamp to hear it
An AI Lab can produce a Transformer, but not GPT
Yang Zhilin breaks ‘breakthroughs happen in industry’ into two layers: first, the inevitable law of field development, as exploratory research gradually shifts to an industrialization process; second, the limits of organizational form. Google Brain is essentially a research organization inserted into a large company; this organizational form can explore new ideas, but it is very hard for it to produce a great system — ‘it might be able to produce a Transformer, but it might not be able to produce GPT.’ He judges that the scientific research system and the education system will later shift their functions, mainly to cultivating talent. This is also how he explains why AGI needs a new organization: technology determines the mode of production, the organization must adapt to the technology, and if it doesn't match, it cannot produce effectively.
— Yang ZhilinRelease yourself from infinite carving
This is the most important lesson he learned at Google. The picture he gives is: there are ten roads in front of you; most people consider how to brake when there is a pedestrian ahead on the road they take, but which of the ten roads to take is the most important thing. At the time, the problem in this field was that the focus was wrong — lowering loss further on a dataset of one or two million tokens, inventing all kinds of bizarre architectures and regularization methods, were all carving techniques; they improved on the dataset but did not see the essence of the problem. The essence is to find first principles: as long as a structure is general enough (all problems can be put into it for modeling) and scalable (put in enough compute and it gets better), you should not do excessive carving at the upper layer. The line he quotes is: if you can solve a problem with scale, don't solve it with a new algorithm.
— Yang ZhilinThe first financing window was only one month
Yang Zhilin describes the timing very concretely: they began concentrating on the first financing round in February, and this window was very short — delay to April and there was basically no chance, but doing it in December or January also had no chance, because at that time there was still covid and people had not yet reacted. In China it exploded in February; in November of the previous year there was still little reaction, so the real window was one month. In the US he did a precise calculation every night: how many flops that corresponded to, how much training cost, how much inference, how many users; the conclusion after calculating was that they had to get at least 100 million USD within a few months. At the time many people thought it might not be possible to raise that much; later it proved possible, and even more. After the financing window came the hiring window in March and April.
— Yang ZhilinHire geniuses first, then raise the floor
Yang Zhilin says hiring is the core of the core; if the people are wrong, even the most superb management skills are useless. There are very few people in the world with real AGI experience, so their early profile was very focused: directly find the matching genius — someone with the ability to operate on models and direct experience training at super-large scale can build it quickly. For a long time the team was thirty or forty people, now eighty, deliberately pursuing talent density and not wanting too many people. His reference is that Google now has several thousand people doing this; if it were cut to one hundred or fifty, it might succeed quickly. Only later did they begin to fill in product, operations, leader-type talent and animals who can push things to the extreme — filling in the floor.
— Yang ZhilinLong context is the memory of the new computer
Why do long context first? Yang Zhilin's analogy is: it is the memory of the new computer; the memory of old computers grew by several orders of magnitude over the past few decades, and the same thing will happen to the new computer. It solves two essential problems — generality and personalization. On generality, when you have a sufficiently lossless long context, in a multimodal architecture you may not even need a tokenizer; you can put the raw signal directly in. On personalization, the core value of AI lands on personalized interaction, and personalization is not achieved through fine-tuning; it is a process defined by a very long context that cannot be replicated. He stresses that this technology has been worked on for more than half a year, not seeing the trend and then gathering two teams to develop quickly; the two are very different.
— Yang ZhilinThe next two milestones: a unified world model and evolution without human data
Yang Zhilin gives two of the most important technical milestones: first, a truly unified world model that can unify different modalities, a truly scalable and general architecture; second, being able to let AI continue to evolve without human data input. His judgment is that many of the reasoning and agent problems discussed today are products after these two problems are solved; there may still be some carving to do, but there will be no fundamental blocker. He also uses this standard to measure product value: if only ten or twenty percent of a product's core value comes from AI, then the thing does not hold; intelligence is always the core incremental value.
— Yang ZhilinOpen source can't catch closed source, because open source itself is still centralized
Yang Zhilin's reason is not how many resources there are, but that the development method has changed: previously everyone could contribute to open source, but now open source is essentially still a centralization that requires the concentration of resources, talent and capital, so closed source is better; it will be a process of consolidation. He adds an angle: if today there were a very leading model still open-sourced, it would most likely be unreasonable. This directly supports his rejection of ‘using open-source models to build a super app’ — without technical capability, finding a super app is impossible.
— Yang ZhilinRushing to find PMF will get you dimensionally reduced
In response to views that only invest in applications, such as ‘10 people can't find PMF, 100 people can't find it either, the key is whether you can find PMF,’ Yang Zhilin's response is that the ceiling of AGI is far higher than everything seen now, and one must think in terms of 10 to 20 years. The historical case he gives is companies that previously did dialogue systems, customer service systems, slot filling; some had decent scale, but all were dimensionally reduced, because no one will use that kind of technology again. And today is still in the process of change — context is getting longer and longer, instruction following is getting stronger and stronger, so the difficulty of customizing a customer service system will get lower and lower and accuracy higher and higher; therefore this is a 100% dimensional reduction. He acknowledges that focusing only on applications has commercially viable opportunities, but the biggest opportunity is not there.
— Yang ZhilinIn their own words · checked verbatim
It might be able to produce a Transformer, but it might not be able to produce GPT.
它可能能产生Transformer 但可能产生不了GPT
Yang Zhilin6:05
If you can solve a problem with scale, don't solve it with a new algorithm.
如果你能用scale解决的问题 你就不要用新的算法去解决它
Yang Zhilin14:13
Only if it is a disruptive thing does it deserve the three letters AGI; otherwise everything we say today is meaningless.
你就是一个颠覆性的东西 它才配得上AGI这三个字 否则我们今天所有说的这些事情都没有什么意义
Yang Zhilin43:36
If you rush to find a PMF, you will very likely find in the end that you have been dimensionally reduced again.
如果是着急的去找一个PMF 你很有可能最后发现 又被降维打击了
Yang Zhilin48:36
Figures
| Moonshot AI team size | Early thirty or forty people, about 80 at the time of the interview | 28:26 |
| Total of Moonshot AI's two financing rounds | Close to 2 billion RMB | 26:25 |
| Financing target Yang Zhilin calculated in the US | At least 100 million USD within a few months | 22:17 |
| First financing window | About one month (February) | 22:17 |
| Price fluctuation of a single machine | 260 one day, 340 the next, then fell back a couple of days later | 30:27 |
| Yang Zhilin's paper citations | More than 22,000 | 1:02 |
Glossary
- scaling law
- As long as you put enough compute into the model, the results get better.
- long context
- The length of text a model can process at one time; Yang Zhilin calls it the memory of the new computer.
- PMF
- A product finding a scenario that is truly needed by the market.
- DiT
- The architecture used by Sora; Yang Zhilin believes it is not yet a general architecture.
- next token prediction
- The only basic principle of language models that has been verified to work.
How to listen
For founders and investors judging the large-model startup window, organizational form and the to-C route, especially those agonizing over whether to chase PMF.
After 1:15, the China-US ecosystem and 2024 predictions are fairly general; you can listen only to the Sora technical judgment part.