The world is too loud. Read what matters.

张小珺·商业访谈录

AI apps haven't broken out — and it's not a product problem, it's a model problem

More than a year after GPT-4, Perplexity is the only AI app from mainstream Silicon Valley VCs to reach 20 million-plus MAU; 90% of the reason is that GPT-4-level capability isn't enough — it can only do information synthesis, not long-horizon reasoning.

AI appsPerplexityLarge modelsGPT-5Business model
Someone long based in the US reviews half a year of the AI app ecosystem, with details on Perplexity's fundraising, retention and moat, plus the physical reason GPT-5 is slow. High information density, specific judgments.

The argument · tap a timestamp to hear it

2:08

A year after GPT-4, AI apps are still boring

The guest's biggest takeaway from the first half of the year: more than a year after GPT-4 came out, AI apps still haven't broken out in a big way, and judged by results it's fairly boring. Setting aside the large-model companies' own apps, among AI apps backed by mainstream Silicon Valley VCs that have reached a meaningful valuation, he can only think of Perplexity — 20 million-plus MAU, 20 million-plus USD ARR, and by year-end possibly 100 million MAU and 100 million USD ARR. Many other companies are still pre-PMF. This judgment is the starting point for the whole episode's discussion: it's not that app teams aren't working hard, it's that the available model capability can't yet support anything bigger.

— Guang Mi
6:11

Information retrieval is the use case that best matches model capability today

The guest divides knowledge workers' creative tasks into three types: combinatorial creativity, exploratory creativity, and transformative creativity. Information retrieval belongs to combinatorial creativity, the capability today's models match best; writing a Tesla investment research report belongs to exploratory creativity, which also requires planning and reasoning; discovering the law of universal gravitation belongs to transformative creativity, and today's large models still struggle to generalize from one example to unseen things. So the complex information Q&A Perplexity does happens to fit existing model capability — a relatively fit PMF. He also mentions education may be the second matching scenario.

— Guang Mi
11:12

AI search only needs to take one or two points from Google to be huge

Perplexity's latest round valued it at 3 billion USD, and the guest explains why it's worth that much: search is still the fattest, biggest market on the internet. Microsoft Bing has only 3.4% of global search share, yet generates 12 billion USD in annual revenue. So if AI search takes just 1% to 2% of Google's share, the business is already big. He also notes that Perplexity's search queries per active user per day are more than three times Google's, and its long-term retention is significantly higher than other AI products; Toutiao's app had roughly 46% next-day retention in 2013, and Perplexity's retention is basically the same.

— Guang Mi
17:50

Perplexity has an 80% chance of being acquired

Asked about Perplexity's endgame, the guest says there's an 80% chance it gets acquired. The reason: startups today all seem to live under the giants' radiation, and this is a core battleground. He cites Meta's glasses integrating a voice version of Perplexity, Siri borrowing a search engine, and Microsoft Bing working at it for ages without earning that good a reputation, while Perplexity's reputation is actually very good. On the moat, he thinks it's the user mindshare that comes from first-mover effect. At the same time he admits Perplexity still has a lot of room to explore on usage frequency, monetization, and how to compete with the giants without colliding head-on.

— Guang Mi
22:50

Apps haven't broken out — 90% is the models not being capable enough

The guest splits the reason AI-native apps haven't seen a systemic breakout into two parts: 90% is that GPT-4-level capability isn't enough — it can only do combinatorial innovation like information synthesis, and can't do long-horizon reasoning or creative work, so everyone still has to grind on the next generation of models, especially reasoning and multimodal capability; the remaining 10% may be a matter of time — based only on GPT-4-level capability, there's still a chance of building a big app in the future. He cites NLP taking twenty years to produce search, and electricity producing the light bulb as its first killer app, and thinks that after more than a year of polishing, people are close to something that feels like PMF — but this needs young product geniuses.

— Guang Mi
36:58

GPT-5 is slow because of GPU construction in the physical world

The guest argues GPT-5 is slower than expected not because of an AI problem, not because there isn't enough data, and not because scaling law has hit a wall, but because of a real, physical-world GPU construction problem. GPT-4 was a several-dozen-fold compute increase over GPT-3; GPT-5 today is only a ten-plus-fold compute increase over GPT-4. H100s only started arriving in bulk in Q4 2023, building clusters takes a lot of time, large GPU clusters are still unstable, and large-scale training only became possible early this year; from getting the GPUs to actually being able to train at scale takes another half year. He expects GPT-5 by year-end, with parameters three to five times larger than GPT-4 and data volume seven or eight to ten times larger.

— Guang Mi
45:35

Large models aren't a good business model — ad platforms are

The guest states plainly that large models are definitely not as good a business model as ad platforms right now. ChatGPT has 100 million-plus DAU; assuming 10% pay and 200 USD a year, that's 2 billion USD — against Google's ad platform revenue of 200 billion-plus USD a year, still just 1% to 2%. An ad platform earns back a new user within 6 to 12 months, and the ROI is calculable; the ROI of buying GPUs for large models can't be calculated, and it combines research attributes with a high failure rate. But he also points out that OpenAI has actually long since earned back its training costs through ChatGPT — the losses are mainly in exploring new model technology, much like a pharma company's new drug R&D.

— Guang Mi
1:01:43

Startups and big companies are a dependency relationship, not a disruption relationship

The guest concludes: startups and big companies are not a disruption relationship but a dependency relationship. First, AI companies still burn too much money; second, the giants are too well positioned. He uses OpenAI as an example: GPUs are constrained by Nvidia, cluster building is constrained by Microsoft, Microsoft is also a 49% major shareholder and backer, Azure is the most important channel to enterprise customers, on the consumer side it can't escape Apple, and in the end it still has to kneel to Apple to get in — while Apple could swap out ChatGPT in a minute. So if OpenAI really wants to disrupt, it seems it can only take a shot at Google. This also explains why the giants are all potentially still beneficiaries.

— Guang Mi

In their own words · checked verbatim

It's been more than a year since GPT-4 came out, and AI apps still haven't broken out in a big way — judged by results, it's fairly boring.

对GP4出来一年多了吧 AI应用还没有大爆发 从结果上看是比较无聊的

Guang Mi2:08

Information retrieval is still the most important use case that matches model capability.

信息检索还是匹配模型能力最重要的Use Case

Guang Mi6:11

Look at startups and big companies — it seems it's not a disruption relationship but a dependency relationship.

对你看创业公司和大公司 好像不是一个颠覆关系 而是一个依赖关系

Guang Mi1:01:43

Figures

Perplexity latest valuation3 billion USD11:12
Microsoft Bing global search share3.4%11:12
Microsoft Bing annual revenue12 billion USD11:12
GPT-4 input price per 1M tokens60 USD35:17
GPT-4o input price per 1M tokens5 USD35:17
GPT-4 output price per 1M tokens120 USD35:17
GPT-4o output price per 1M tokens15 USD35:17
ChatGPT DAU100 million-plus45:35

Glossary

PMF / Product-Market Fit
The stage where a product matches market demand and can grow organically.
RAG / Retrieval-Augmented Generation
Retrieving external knowledge first, then having the model generate an answer, to make up for gaps in the model's own knowledge.
MOE / Mixture of Experts
A large-model architecture made up of multiple expert subnetworks, activating only part of the parameters each time.
Scaling Law
The empirical rule that model capability improves as parameters, data and compute scale up.
Latency
The waiting time between sending a request and receiving the model's response.

How to listen

Who it's for

VCs watching AI app investment and startups, AI product leads hunting for PMF, and technical and strategy people who want to understand why GPT-5 is slow and whether the large-model business model holds up.

Skip

After 1:03:32 the China-US innovation ecosystem comparison is fairly generic and can be skipped.