AI apps haven't broken out — and it's not a product problem, it's a model problem
More than a year after GPT-4, Perplexity is the only AI app from mainstream Silicon Valley VCs to reach 20 million-plus MAU; 90% of the reason is that GPT-4-level capability isn't enough — it can only do information synthesis, not long-horizon reasoning.
The video won't play here. Listen to the audio instead:
The argument · tap a timestamp to hear it
A year after GPT-4, AI apps are still boring
The guest's biggest takeaway from the first half of the year: more than a year after GPT-4 came out, AI apps still haven't broken out in a big way, and judged by results it's fairly boring. Setting aside the large-model companies' own apps, among AI apps backed by mainstream Silicon Valley VCs that have reached a meaningful valuation, he can only think of Perplexity — 20 million-plus MAU, 20 million-plus USD ARR, and by year-end possibly 100 million MAU and 100 million USD ARR. Many other companies are still pre-PMF. This judgment is the starting point for the whole episode's discussion: it's not that app teams aren't working hard, it's that the available model capability can't yet support anything bigger.
— Guang MiInformation retrieval is the use case that best matches model capability today
The guest divides knowledge workers' creative tasks into three types: combinatorial creativity, exploratory creativity, and transformative creativity. Information retrieval belongs to combinatorial creativity, the capability today's models match best; writing a Tesla investment research report belongs to exploratory creativity, which also requires planning and reasoning; discovering the law of universal gravitation belongs to transformative creativity, and today's large models still struggle to generalize from one example to unseen things. So the complex information Q&A Perplexity does happens to fit existing model capability — a relatively fit PMF. He also mentions education may be the second matching scenario.
— Guang MiAI search only needs to take one or two points from Google to be huge
Perplexity's latest round valued it at 3 billion USD, and the guest explains why it's worth that much: search is still the fattest, biggest market on the internet. Microsoft Bing has only 3.4% of global search share, yet generates 12 billion USD in annual revenue. So if AI search takes just 1% to 2% of Google's share, the business is already big. He also notes that Perplexity's search queries per active user per day are more than three times Google's, and its long-term retention is significantly higher than other AI products; Toutiao's app had roughly 46% next-day retention in 2013, and Perplexity's retention is basically the same.
— Guang MiPerplexity has an 80% chance of being acquired
Asked about Perplexity's endgame, the guest says there's an 80% chance it gets acquired. The reason: startups today all seem to live under the giants' radiation, and this is a core battleground. He cites Meta's glasses integrating a voice version of Perplexity, Siri borrowing a search engine, and Microsoft Bing working at it for ages without earning that good a reputation, while Perplexity's reputation is actually very good. On the moat, he thinks it's the user mindshare that comes from first-mover effect. At the same time he admits Perplexity still has a lot of room to explore on usage frequency, monetization, and how to compete with the giants without colliding head-on.
— Guang MiApps haven't broken out — 90% is the models not being capable enough
The guest splits the reason AI-native apps haven't seen a systemic breakout into two parts: 90% is that GPT-4-level capability isn't enough — it can only do combinatorial innovation like information synthesis, and can't do long-horizon reasoning or creative work, so everyone still has to grind on the next generation of models, especially reasoning and multimodal capability; the remaining 10% may be a matter of time — based only on GPT-4-level capability, there's still a chance of building a big app in the future. He cites NLP taking twenty years to produce search, and electricity producing the light bulb as its first killer app, and thinks that after more than a year of polishing, people are close to something that feels like PMF — but this needs young product geniuses.
— Guang MiGPT-5 is slow because of GPU construction in the physical world
The guest argues GPT-5 is slower than expected not because of an AI problem, not because there isn't enough data, and not because scaling law has hit a wall, but because of a real, physical-world GPU construction problem. GPT-4 was a several-dozen-fold compute increase over GPT-3; GPT-5 today is only a ten-plus-fold compute increase over GPT-4. H100s only started arriving in bulk in Q4 2023, building clusters takes a lot of time, large GPU clusters are still unstable, and large-scale training only became possible early this year; from getting the GPUs to actually being able to train at scale takes another half year. He expects GPT-5 by year-end, with parameters three to five times larger than GPT-4 and data volume seven or eight to ten times larger.
— Guang MiLarge models aren't a good business model — ad platforms are
The guest states plainly that large models are definitely not as good a business model as ad platforms right now. ChatGPT has 100 million-plus DAU; assuming 10% pay and 200 USD a year, that's 2 billion USD — against Google's ad platform revenue of 200 billion-plus USD a year, still just 1% to 2%. An ad platform earns back a new user within 6 to 12 months, and the ROI is calculable; the ROI of buying GPUs for large models can't be calculated, and it combines research attributes with a high failure rate. But he also points out that OpenAI has actually long since earned back its training costs through ChatGPT — the losses are mainly in exploring new model technology, much like a pharma company's new drug R&D.
— Guang MiStartups and big companies are a dependency relationship, not a disruption relationship
The guest concludes: startups and big companies are not a disruption relationship but a dependency relationship. First, AI companies still burn too much money; second, the giants are too well positioned. He uses OpenAI as an example: GPUs are constrained by Nvidia, cluster building is constrained by Microsoft, Microsoft is also a 49% major shareholder and backer, Azure is the most important channel to enterprise customers, on the consumer side it can't escape Apple, and in the end it still has to kneel to Apple to get in — while Apple could swap out ChatGPT in a minute. So if OpenAI really wants to disrupt, it seems it can only take a shot at Google. This also explains why the giants are all potentially still beneficiaries.
— Guang MiIn their own words · checked verbatim
It's been more than a year since GPT-4 came out, and AI apps still haven't broken out in a big way — judged by results, it's fairly boring.
对GP4出来一年多了吧 AI应用还没有大爆发 从结果上看是比较无聊的
Guang Mi2:08
Information retrieval is still the most important use case that matches model capability.
信息检索还是匹配模型能力最重要的Use Case
Guang Mi6:11
Look at startups and big companies — it seems it's not a disruption relationship but a dependency relationship.
对你看创业公司和大公司 好像不是一个颠覆关系 而是一个依赖关系
Guang Mi1:01:43
Figures
| Perplexity latest valuation | 3 billion USD | 11:12 |
| Microsoft Bing global search share | 3.4% | 11:12 |
| Microsoft Bing annual revenue | 12 billion USD | 11:12 |
| GPT-4 input price per 1M tokens | 60 USD | 35:17 |
| GPT-4o input price per 1M tokens | 5 USD | 35:17 |
| GPT-4 output price per 1M tokens | 120 USD | 35:17 |
| GPT-4o output price per 1M tokens | 15 USD | 35:17 |
| ChatGPT DAU | 100 million-plus | 45:35 |
Glossary
- PMF / Product-Market Fit
- The stage where a product matches market demand and can grow organically.
- RAG / Retrieval-Augmented Generation
- Retrieving external knowledge first, then having the model generate an answer, to make up for gaps in the model's own knowledge.
- MOE / Mixture of Experts
- A large-model architecture made up of multiple expert subnetworks, activating only part of the parameters each time.
- Scaling Law
- The empirical rule that model capability improves as parameters, data and compute scale up.
- Latency
- The waiting time between sending a request and receiving the model's response.
How to listen
VCs watching AI app investment and startups, AI product leads hunting for PMF, and technical and strategy people who want to understand why GPT-5 is slow and whether the large-model business model holds up.
After 1:03:32 the China-US innovation ecosystem comparison is fairly generic and can be skipped.