Define yourself as an LLM company and none of you survive
Li Di says it isn't that nine of ten LLM companies die — it's that if you define yourself as an LLM company, not one survives. Because the value you create and the money you collect are two different things: API pricing hugs cost, and the profit ceiling is 0.2 cents per thousand tokens.
The video won't play here. Listen to the audio instead:
The argument · tap a timestamp to hear it
API pricing hugs cost, so the profit ceiling is fixed
Li Di dismantles the commercial fantasy of large models with a single calculation: a friend spends $20 a month on GPT, but says outright that the moment a cheaper or better service appears he'll switch, with no stickiness at all. It is a pure relationship of interest, and price is judged against cost. At 0.2 cents per 1000 tokens, the only way to earn more is to push cost down, and if cost won't come down you have to cut quality — which is exactly where OpenAI's recent quality decline comes from. So the profit ceiling is locked at 0.2 cents. He adds a line from his mother: saving money is not earning money — earning money is unlimited, while saving money can only come out of the money you already have.
— Li DiMonopolise all media and you still make only a few hundred thousand a month
A peer came to Li Di saying he wanted to use GPT to write reports, fine-tune them into his own outlet's style, and replace editors entirely. Li Di asked what he'd be willing to pay; the answer was 0.2 cents per 1000 tokens. That works out to a few cents an article, and 200-plus articles a month comes to a few tens of yuan. The editors being replaced were paid 1500 yuan and up per article. Li Di asked why he couldn't offer a percentage discount; the man said no — you have competitors, the price is what it is. The conclusion: even if you monopolised all the content of every media outlet in China, creating enormous value, in the end you'd make only a few hundred thousand a month. Value is not revenue; revenue is what counts as value.
— Li DiYou can't sell just a model, you have to sell a whole person
Li Di says he made two points last year: defining yourself as an LLM company leaves you essentially no way to survive; and a model is not better the bigger it gets — the bigger it gets, the less differentiated and more commoditised it becomes. Xiaoice's approach is to go for high added value — "you can't provide only the inner technical layer, you can't provide a whole, a person has many capabilities, but you can't personify a person into a set of capabilities." He cites SenseTime's face recognition: face recognition matters enormously, it touches every part of life, yet it earns nothing, so the company was forced into hardware-plus-software, making its money on the 2000 yuan of a security door — that isn't AI money, but it gets booked as AI revenue. That kind of business cannot last.
— Li DiA 3.5B model runs on a T4; a large model's latency is unusable
Xiaoice open-sourced in Japan, and in stability.ai's evaluation the top-ranked general-purpose model was Xiaoice's, with Meta's Llama 65b (65 billion parameters) second — while the first-place Xiaoice model has only 3.5b. The reason is focus. The engineering math is clear: 3.5b runs on a T4, around 7b runs on a V100, and anything under 13b runs on a single A100 without distributed setup, so latency stays very smooth. A model as large as GPT, unless you spend a great deal on an array, has high latency — unusable on smart devices, where saying something and waiting 5 seconds for a reply is unacceptable. So large models should stay in the lab exploring new knowledge, training other models, and not leave the house.
— Li DiPeople think the technical paradigm is settled; it has barely begun
Li Di says one of the Chinese market's misreadings of large models is taking OpenAI's success to be brute-force aesthetics and trying to copy it — but they make a big mistake: they think this technical paradigm is already determined, when in fact it has only just begun. The underlying principles aren't even clear, so on what basis do you assume today's leader stays the leader? He offers an analogy: this is a 1000-metre race and we've run two steps, two metres, and the one whose head is a little further forward may be closer to being wrong. Technological development is cyclical — a breakthrough is found, innovation emerges, a bottleneck is hit, a few years or a few months of struggle, then someone breaks through again — it is not exponential.
— Li DiChain of thought at 65B is an observation, not a conclusion
Google observed at the time that models below 65B had no obvious chain-of-thought ability and that above 65B it suddenly appeared, so it wrote a paper saying large models' chain-of-thought ability exists only at 65B. Li Di says that is a record of an observation, not a scientific conclusion, but everyone treated it as a scientific conclusion. Today models of a few B have chain-of-thought ability. He compares it to Mendel studying the laws of heredity: that too was observation without underlying principle, and it couldn't be treated as a conclusion. This is also why everyone believed at the time that bigger models were better.
— Li DiLarge models are the number two player's offensive weapon
Li Di says another kind of company isn't an LLM company but an existing business plus a large model, used to compete. When Baidu launched Ernie Bot many people said it was going to do search; Li Di judged that impossible — it had to be added onto some number two, three or four business, and indeed it turned out to be Baidu Cloud. This is the number two's offensive weapon; if you're the number one and you deploy it, isn't that suicide — raising your own costs and pushing your profit down. Xiaoice was the same: at the time it had huge conversational traffic in China and no incentive to run the most expensive model, since a single interaction that cost 0.5 li would become several cents and the bill would drown you instantly.
— Li DiPublishing a paper is leaking a trade secret
The Xiaoice team avoids publishing papers wherever possible. Li Di gives three reasons. First, publishing means making your method public; in the US publishing a paper is much like filing a patent, but in China someone can follow your paper and file a patent on it, so every time they publish they find a way to hide the real thing. Second, once a researcher takes publishing as the achievement, they go into fields that are easy to publish in and avoid fields that are hard to publish in but may matter, and papers, in order to be attractive, often mean the engineering isn't rigorous enough — having done 0 to 1, they're already looking for the next 0, like a bear dropping one corn cob for the next. Third, papers follow the person, not the company: once someone is famous they get poached, and labour costs rise.
— Li DiIn their own words · checked verbatim
It's not that of 10 large-model companies, 9 die, or of 48, 2 survive. If it's further defined as a large-model company, not one of them breaks through. No one can run a model company, unless the model company's business model changes — not the API Code way.
大模型公司不是10个 死9个或者48个活2个 如果它再界定到大模型公司 是一个都破不了 没有人能做一个模型公司 除非模型公司 商业模式发生变化 不是API Code的方式
Li Di0:00
This is a relationship of interest. I can use it; whoever is better, I use. And the price is judged against cost.
这是一种利益关系 我能用 谁好我就用谁 而且价格是以成本来考量的
Li Di19:16
Earning money is unlimited. Saving money can only come out of the money you have. At most you eat and drink nothing at all, and you can only save what you earn.
挣钱是无限的 省钱只能从你有的钱里面省 你最多完全不吃不喝 你也就只能省下你的收入
Li Di20:17
This is a 1000-metre race and we've run two steps, two metres. Its head is a little further forward; it may be closer to being wrong.
这是一千米才跑了两步 才跑了两米 它头往前了一点 它可能离错误更近了
Li Di32:21
This is a record of an observation. It is not a scientific conclusion. Everyone just took it to be a scientific conclusion.
这是一个观测 结果的记录 它不是一个 科学的 结论 大家就认为这是一个科学结论
Li Di35:24
This is the number two's tool. This is the number two's offensive weapon. But if you're the number one and you deploy it, aren't you committing suicide?
这是老二的工具 这是老二的攻击性武器 但你自己要是老大你就上这个 你不是自杀吗
Li Di39:26
I hope it can survive, survive forever. But the way it survives is by permeating people's lives, by forming countless intricate long-wall relationships with a great many people, so that no one can shut it down.
我希望他能活下去 永远活下去 但他活下去的方法 就是他渗透到人们的生活中 跟大量的人产生千丝万缕的长城的关系 就没有人能够停掉他
Li Di1:20:43
Figures
| GPT API pricing | 0.2 cents per 1000 tokens | 19:16 |
| Media editor pay comparison | An editor gets 1500 yuan and up per article; model-generated output at 200-plus articles a month is worth only a few tens of yuan | 20:17 |
| Large-model parameter threshold | Above 65B means 65 billion parameters | 25:18 |
| Change in Xiaoice interaction cost | Originally 0.5 li per interaction; after running the most expensive model it became several cents | 39:26 |
| Microsoft share price | 30 when Li Di joined, 300 at the spin-off | 1:10:36 |
| Share of Xiaoice systems that were bullied | 20% | 1:18:41 |
| Number of LLM companies in China | 76, as heard last month | 57:27 |
Glossary
- API
- A way of selling model services billed by call volume; Li Di argues this model prices at cost and can't be made into a business.
- chain of thought
- A model's ability to reason step by step; Google once observed it appearing only above 65B, which Li Di says was an observation, not a conclusion.
- mixed model
- The route Xiaoice advocates: mixing large, medium and small parameter-scale models rather than having one enormous model handle everything.
- digital employee
- Xiaoice's 2B business direction: using AI to fill roles inside a company that directly generate revenue, rather than cost-centre roles.
How to listen
Founders and investors building LLM startups who are torn between piling on more parameters and pivoting to applications, plus AI engineers who want the small-model engineering math.
The rapid-fire Q&A after 1:22:48 (books, idols, movies) is low-density and skippable.