The world is too loud. Read what matters.

张小珺·商业访谈录

Li Zhifei Talks People Out of General-Purpose Large Models: China Will Not Have a Second OpenAI

A month ago he said he wanted to be China's OpenAI; now he says that proposition most likely does not exist — because consensus formed too fast, the difficulty was rapidly lowered by open source and compute, and it is not even clear OpenAI itself can win commercially in the end.

Large modelsStartupsOrg structureAGIBusiness model

The video won't play here. Listen to the audio instead:

The value of this episode is not technical education but a practitioner publicly overturning his own judgment from a month earlier, and giving concrete reasons to talk people out of general-purpose large models.

The argument · tap a timestamp to hear it

10:45

ChatGPT's real innovation is "one unified system"

Li Zhifei reduces the biggest difference between this wave and the past decade to "general": before, speech recognition was one team and machine translation another, and inside Google they were completely different teams, each with its own training data and its own code system; in the future a large model may use one unified system to do speech, images, translation, chat and protein structure prediction at the same time. He agrees with Microsoft's article saying GPT-4 is a spark of AGI, on the grounds that the things general intelligence requires — generality, high abstraction rather than memorizing surface patterns, planning ability, and freely combining tasks to do new ones — GPT-4 already has.

— Li Zhifei
14:00

Self-supervision means the internet gives feedback at every step

He gives an explanation more accurate than "throw the kid in the sea to learn swimming": predict the next word based on the preceding words. We are chatting, and after "chat" comes "ting"; the internet has massive amounts of text like this, and the AI is rewarded when it predicts right and punished when it predicts wrong. It is called self-supervised because this data was not labeled — unlike machine translation, which requires a Chinese sentence with an English annotation, or speech recognition, which requires audio with a text annotation — instead the sequence text is taken directly from the internet, and every time it predicts the next word the data itself gives clear feedback.

— Li Zhifei
22:40

Google's problem is organizational form, not capability

He judges that on capability Google could crush OpenAI: the number of people who can get things done may be ten times OpenAI's, and compute and data can to some degree crush it too. But its organizational form may not be able to compete — the research division is completely separate from product, and mobilizing data and resources from YouTube, Search and Cloud, or even shipping a product, is a cross-departmental ordeal; inside a department there are too many smart people, each with their own methodology, some wanting to use GPT, some BERT, some T5. OpenAI, by contrast, has belief and is product-driven in its iterations, and is extremely focused and unified in methodology, faith and execution. His exact words: an opponent ten times stronger will not necessarily beat you on something this uncertain.

— Li Zhifei
26:53

Consensus formed too fast and washed the moat away

His judgment a month ago was that the moat was extremely high, early investment enormous, and very few people truly willing to do it. In the past month or two everything changed: first, there will be many players, more than a dozen, because this was too quickly made a consensus — the most important thing of the next ten or twenty years; second, the difficulty depends on how you do it. If you keep pushing the capability ceiling like OpenAI or Google, of course it is extremely hard, but if you build around your own application scenarios the difficulty drops sharply, because open-source models, Nvidia's stronger computing platforms, and countless people's interpretations and papers are all pulling down the thresholds at the three levels of compute, algorithms and data. So he no longer thinks you must start by building a separate company, finding the strongest people, and going into seclusion for twelve months.

— Li Zhifei
34:58

China does not have an OpenAI-style organization

He says directly, "whether China necessarily has an organization like America's OpenAI, I think most likely it does not exist," and calls wanting to be China's OpenAI a false proposition. He accepts that China needs many large models, but is not sure it has the ability to build the kind that explores the capability ceiling, and stresses that building large models is not a single path. To Robin's remark that "China does not need a second large model," his response is: China certainly needs many large models, but whether it can build one as formidable as OpenAI, exploring the ceiling, he is not sure. He also says that even globally, OpenAI is not necessarily the one laughing last.

— Li Zhifei
43:36

This generation of AI companies also has ten times the supply of the last

He sums up the shared challenge of the previous generation of AI companies as a bad business model: high investment, low output, every company in a very poor state of commercialization. The good side of this generation is that application scenarios will certainly far exceed the last, and demand may be ten or a hundred times greater; the bad side is that supply may also be ten times the last generation's, which will make many AI companies suffer just like the last generation. When he did AI in 2012, in China the companies doing large-scale speech recognition and voice assistants were few and far between, let alone second-tier internet and traditional companies; today, after large models appeared, the consensus far exceeds that of those years, and a consensus that strong, reached in such a short time, is good for the industry and society, but for players in the game it means investment and competition will be brutally fierce.

— Li Zhifei
1:03:42

The discouragement: don't go in big and loud, figure out deployment first

He says he hopes to talk some people out of doing general-purpose large models, and stresses this has nothing to do with personal competition, because he has no conflict with them at all. There are two reasons: it is very hard, and commercially the competition will be extremely fierce — if what you are building is a very general large model but you have not carefully thought about the best scenario to deploy it in and how the business model works, you will suffer greatly later. His advice: first, do not go in big and loud, spending and hiring at any cost and without regard for cost; second, think as much as possible about deployment and the business model. He also says OpenAI would be suffering too without Microsoft, and even today he remains quite pessimistic about OpenAI's business model — relative to its investment, whether it can laugh last commercially is really hard to say.

— Li Zhifei
1:06:23

Among the giants he bets on ByteDance, because Zhang Yiming reads papers himself

Asked about the giants' landscape, he says ByteDance may be the hope of the whole village, because "they get it — he reads papers himself, talks to people himself every day, finds engineers," and names Zhang Yiming as extremely strong in execution, China's strongest player at brute-forcing miracles. On Alibaba, his judgment is that it must do it, because cloud services will have to have such a large model to empower partners, otherwise it is hard to say cloud services are competitive. He goes further: for the giants, large models are standard equipment, only a matter of share size; and it is fairly certain that a year or half a year from now, if you have no model or no move in large models, the capital markets will simply stop looking at you.

— Li Zhifei

In their own words · checked verbatim

Whether China necessarily has an organization like America's OpenAI, I think most likely it does not exist.

中国是不是一定存在一个跟美国OpenAI一样的这种组织,我觉得大概率是不存在的

Li Zhifei34:58

Demand may be ten or a hundred times greater, that is certainly good, but the bad part is that I think supply may also be ten times the last generation's.

需求可能是以前的十倍百倍,这个肯定是好的,但是坏的地方呢,我觉得供给可能也是上一代的十倍

Li Zhifei43:36

If what you are doing now is a very general large model, but you have not carefully thought about the best scenario to deploy it in and how your business model works, you will suffer greatly later.

如果说你现在去做的,就是一个非常通用的大模型,但是你没有仔细想过,我最好落地在什么场景下,我的商业模式怎么做的话,你后面会非常痛苦

Li Zhifei1:03:42

If, a year or half a year from now, you have no model or no move in large models, then the capital markets will simply stop looking at you.

如果说,过一年,或者半年以后,你在大模型上,没有模型,或者没有动作,那资本市场就是,完全不看你了

Li Zhifei1:06:23

Figures

His estimate of the lead time of large models over opponents like Googleat least six months, eight months or more22:19
Parameter scale of Mobvoi's large modeltens of billions1:05:40
His projection of the number of companies in China building large models two years outmore than 5044:28
His conclusion on the intensity of competition in ChinaChina is ten times the competition of the US (roughly twice the startup supply × roughly one-fifth the customer price)1:01:36
How long Mobvoi kept building its large model in 2020stopped after about 8 months29:23

Glossary

self-supervised
Not relying on human labeling, but predicting the next word from sequence text, with the data itself giving feedback at every step.
in-context learning
The ability of a model to learn a new task from examples in the prompt alone, without changing its parameters.
AGI
A single system that can generally perform many tasks, rather than training a separate model for each task.

How to listen

Who it's for

Founders assessing whether to build large models, investors watching the AI track, and anyone who wants to understand the judgment that "consensus forming too fast destroys moats."

Skip

1:47–9:06, the history of his NLP education and technology iterations — personal background setup.