Large Models Are the Big Players' Game; Startups Shouldn't Follow Jasper's Path
Large models are not a replacement for search engines but a supplement. The real opportunities lie in vertical domains like healthcare and materials where large models can't capture the data, while general-purpose large models belong only to the giants.
The video won't play here. Listen to the audio instead:
The argument · tap a timestamp to hear it
Technology only explodes when end users can perceive it
Mengqiu said that in the previous wave of computer vision AI, they didn't look at it at all, because no matter how advanced the technology, if it's only 2B, end users don't feel it, and it's hard to explode. All technology and underlying platforms only explode when mid-end users can perceive them. Last September, when he saw the effects of Stable Diffusion and Midjourney, he told his team, "We need to look at AI," because this is something end users can perceive. The key to ChatGPT's explosion is not Chat or GPT, but that it made a technological change easy to use, interesting, fresh, and worth using repeatedly.
— Meng QiuThe key to the iPhone moment is product definition capability
Mengqiu believes that when the iPhone was launched, almost no technology was originally created by Apple. Touch screens, multi-touch, these building blocks already existed. Nokia didn't lack them or couldn't use them, but Apple had product definition capability, assembling all components into a mobile computer. Similarly for ChatGPT, the underlying is GPT, but there must be a good product carrier to make it perceptible to end users, so that future application scenarios and commercial imagination space are large enough. He also mentioned that OpenAI's female CTO Mira Murati played a big role in pushing the release of ChatGPT.
— Meng QiuThe previous AI wave couldn't rise due to belief and resource issues
Mengqiu analyzed that the previous AI wave was essentially about using deep learning and neural networks for AI under limited computing power and without much reinforcement learning. It worked well in video and audio recognition, but no one thought, "If I had unlimited computing power, what could I make of it?" This is a matter of technical route and belief. DeepMind is the strangest institution in the world, burning hundreds of millions of dollars a year. One of its three founders focused on commercialization, but in the end, it was sold to Google. Doing AGI is something some believe in and some don't. Those who believe are also limited by funding conditions. Who would have long-term money to support you burning so much?
— Meng QiuTwo key nodes in OpenAI's success
Mengqiu believes OpenAI started as a public welfare organization composed of a group of top-tier figures. In the early years, they burned a lot of money, and many people withdrew. The first key point was Sam Altman bringing in Microsoft investment. Microsoft made an extremely complex convertible bond deal structure with them. In those years, OpenAI didn't need to consider revenue; Microsoft provided infinite computing power and specially built supercomputing centers. The second key point was that they themselves were a bunch of researchers doing AGI, with very fast thinking, oriented towards application scenarios and products, and they launched ChatGPT before pushing GPT-4. Mengqiu said this equity design is hard to replicate.
— Meng QiuInvestors look fanatical, but they act very rationally
Mengqiu observed that this round of AI wave differs from the metaverse, digital humans, blockchain, and VR bubbles in that everyone feels this thing has long-term value support, so they are actively looking. But are they actively investing? Not really. Especially top-tier funds, their strategy might be "I'll invest in everything, I don't know who will succeed," investing in large models, application layers, middle layers, and computing architectures. Most funds are still on the sidelines because things change a lot; yesterday's understanding of this matter will change today.
— Meng QiuChina's SaaS implementation difficulty is due to uneven ground
Mengqiu said that in the past, China's SaaS couldn't be implemented not because people didn't do SaaS well, but because the ground is uneven—many enterprises haven't even achieved basic processes, let alone data. What would you use SaaS or AI for them? They don't have data. This is very different from American enterprises, where most have a high degree of digitization, and many startups are very digitalized from day one. So there are few customers willing to pay you.
— Meng QiuDon't follow Jasper's old path
Mengqiu said that it's easy to cross out things they think shouldn't be invested in, but there's no answer to what should be invested in. What's crossed out is what large models can easily do: building vertical domain models on top, or fine-tuning based on APIs—these have little value. He gave the example of Jasper.AI. Jasper was originally an early deep partner of OpenAI, calling their API to generate marketing copy with good results. But as soon as ChatGPT came out, it was done, because the needs of a large number of small and medium customers could be directly met by ChatGPT. So the criterion is: is the data in your vertical domain niche enough that large models can't capture it, and can the capabilities you need be covered by large models?
— Meng QiuThe high barriers for this round of startups are talent density and funding
Mengqiu summarized the characteristics of this round of startups: First, talent density requirements are very high. OpenAI doesn't have many people, and Google's internal AI research institute plus DeepMind doesn't have many either. It's definitely not about piling up manpower, but each person has a particularly luxurious background, and these people are very expensive. Second, the funding threshold is high. A rumored company burning 50 million to 100 million USD a year is normal, because when algorithms and architectures are not good, running a model is very expensive. Third, commercialization implementation is relatively long. Fourth, it requires at least a team of algorithm, architecture, and engineering people. That's why everyone is forming groups and banding together.
— Meng QiuLarge models are overall the big players' game
Mengqiu said that large models are still the big players' game, and he doesn't think a new giant will suddenly appear to challenge them. OpenAI's emergence first had a natural giant demeanor, originally a bunch of very impressive people; second, before it could achieve such effects, the giants were dozing off. Now the first wave of awakening in China is the giants themselves. But he believes AI will reshape the workflows of all industries, and talent models and efficiency models will all be completely changed. He also said that if China becomes a closed market in the future, 1.4 billion people are big enough. From a commercial competition perspective, there must be at least two players. In China, there is nothing that fewer than three players are doing.
— Meng QiuIn their own words · checked verbatim
This is why I've always felt that large models are still the big players' game, no matter which big player. I don't think a new giant will suddenly appear within China to challenge them.
这就是为什么我一直觉得大模型仍然是大厂的菜,不管是哪个大厂吧,就我不会觉得有突然又出现一个新的,在中国境内啊,出现一个巨头出来挑战。
Meng Qiu48:38
It doesn't pursue the accuracy of the answer; it pursues the answer. So they have a random factor in the model, and its model always predicts the next word.
它不追求答案的准确,它追求的是答案的,所以他们在模型中有一个随机因子的,就它的模型就永远预测下一个词是什么。
Meng Qiu53:31
Figures
| OpenAI annual burn rate | 50 million to 100 million USD | 36:48 |
| Funding threshold for large model startups | Over 200 million USD | 57:32 |
| Funding threshold for mobile internet era startups | 1 million USD | 58:33 |
| Funding threshold for consumer startups | 1 to 2 million USD | 58:33 |
| Google's founding time | 1997 to 1998 | 41:22 |
| When Google started working on its business model | Around 2001 | 41:22 |
| Interval between search engine business model and founding | Three to four years | 41:22 |
Glossary
- Killing App
- An application that both has a large number of users and generates new commercial behaviors.
- Prompt Engineering
- Guiding generative AI to produce more desired results by constructing different prompts.
- COT
- Chain of Thought, a method to guide the model to reason step by step.
- Fine tuning
- Continuing to train an existing large model with specific data to adapt to vertical scenarios.
- Middle layer
- A tool or platform layer between the underlying large model and the application layer.
How to listen
Entrepreneurs and investors focused on large model startups and primary market investment, as well as engineers who want to understand the real temperature of this AI wave.
After 1:07:35, the part recommending public accounts and papers can be skipped.