Li Xiang: No general agent within five years — what comes first is an agentOS
Li Xiang's judgment: a general agent will not appear within five years; what will actually appear is an agentOS, on which each profession builds its own professional agent. Because a production tool has to be able to act and replace professional work, not just give advice.
The video won't play here. Listen to the audio instead:
The argument · tap a timestamp to hear it
AI has three tiers, and only production tools are worth money
Li Xiang sorts AI products into three tiers: information tools, assistance tools, production tools. Information tools only give reference; assistance tools raise the competitiveness of an existing product; only production tools truly replace professional work and cut working hours. His test is hard: are you willing to pay for it. Right now the only entry-level production tools his colleagues around him pay for out of their own pockets are two — Cursor for the programming colleagues, and OpenAI's Deep Research for the business analysis and strategy team. He holds that the most important criterion for an agent is whether it is a production tool, whether it can genuinely replace me in doing professional work.
— Li XiangWhat DeepSeek teaches is the four steps of human best practice
What Li Xiang learned from DeepSeek is that it used human best practice in an extremely minimal way. When building capability, four steps: first do research, then do development, then express the capability, and finally turn it into business value. When doing business, four steps: first do index analysis, then set the goal, then make a strategy and execute it, and finally do reflection and review. He says humans often forget best practice as they go — when they hit a problem they only want to change the strategy, without doing a review, without doing market analysis, without re-setting the goal. And strictly following best practice is anti-human, so a good organisation, and an outstanding person, often has to fight against human nature.
— Li XiangLiang Wenfeng is self-disciplined — that means holding to best practice
Li Xiang talked with Liang Wenfeng once last September, and two things struck him most: first, he is an especially self-disciplined person, and the biggest mark of that self-discipline is being able to hold to what he believes in, to hold to best practice, to fight against the laziness and shortcut-taking in human nature; second, he researches and learns best practice and the best methodologies worldwide. Li Xiang also revealed that DeepSeek's open source accelerated Li Auto's VLA work by nine months and saved roughly several hundred million in cost, so Li Auto open-sourced its operating system too — purely to thank DeepSeek, not as company strategy.
— Li XiangVLA training has three steps, starting with the VL base
Li Xiang walks through VLA training step by step. Step one, train the VL base; the current version is a 32B cloud model, and the data has three parts: 3D Vision, high-definition 2D Vision (3 to 5 times sharper than open-source VLMs), and corpus that joins images with semantics — for example putting the navigation map together with human understanding of the map, data only Li Auto has. After the base is trained it is distilled into a 3.2B on-device MOE model made up of eight experts, because dual Orin X and Thor U running the full 3.2B model directly cannot hit the frame rate. Step two, post-training, bringing action in, like going to a driving school to learn to drive; the model expands from 3.2B to close to 4B, COT does only two to three steps, and it also does diffusion to predict the trajectory for the next 4 to 8 seconds. Step three, reinforcement: first RLHF to align with humans, then pure RL using data generated by the world model, scored on three items — comfort, collision, traffic rules.
— Li XiangThe world model has three stages, ending as an L4 operating system
Li Xiang distinguishes two readings of the world model: in robotics, diffusion predicting the next few seconds after action is treated as the world model; Li Auto classifies that as part of VLA's capability, and the real world model is a reconstructed plus generated traffic physics world. It has three stages: the first stage is used for exams, the second generates training data, the third becomes the operating system for future L4-level autonomous vehicles. Because you cannot write a traditional IT software to operate a self-driving car running on the road with nobody in it. He also gives numbers: verification cost per ten thousand kilometres falls from 170,000 to 180,000 RMB to just over 4,000, essentially all compute cost, and it can reproduce real scenes 100%.
— Li XiangETC uses a rule algorithm, solving a three-to-four-month problem in a week
Li Xiang uses ETC as an example of why he does not give up on tools. VLM is terrible at judging position; with two or three ETC lanes it can still judge left, middle, right, but with a dozen ETC lanes like on the airport expressway it gets confused. The team kept feeding the VLM more corpus without solving it, because this is an architectural problem of VLM. Li Xiang said, why can't ETC be solved with a rule algorithm — at most 15 lanes, write a program and it takes a day, or even three days. The team solved the problem quickly, ETC became very stable, and in under a week they solved what three or four months had not solved, what costly approaches had not solved. His principle: what is deterministic and can be solved with rules means lower energy consumption, lower compute consumption and higher accuracy.
— Li XiangL4 is stuck on on-device compute, not on algorithms
Li Xiang judges that the ceiling of L3 and L4 is set by model scale. Today the car carries a 3.2B MOE, roughly 4B counting action; two Orin X or one Thor U have 64G of memory and can run a model of up to 30B, but the frame rate falls short. To meet the frame rate for traffic or robotics, model scale has to be pushed down, and on-device model scale is constrained. If one day a 32B model could be put directly on-device, L4 might be achieved, and the cloud model would also expand to 320B. He says this year's compute is basically at the L3 level, and the earliest in Q3, the latest in Q4, we will see products with real L3 capability, but regulations and so on still have to be resolved.
— Li XiangThree to seven people form a stronger brain and heart
Li Xiang describes the way of building energy he worked out in his review: three people are enough to form effective support; three people do not turn inward, they face outward together, and through arguing and thinking they form a more comprehensive and stronger brain, and once a judgment forms they also form a stronger heart — when you lack energy, the others replenish you. He says generally three to seven people is the most stable; fewer than three is too few, more than seven is too many, so today when designing many structures he deliberately designs three-to-seven-person combinations as the mental and psychological support structure. He also says that when energy exists between people, arguments, discussions and fights are a more complete brain; when that energy disappears, these are internal friction.
— Li XiangIn their own words · checked verbatim
I think only when artificial intelligence becomes a production tool does the real explosion of artificial intelligence arrive.
我觉得人工智能变成生产工具 然后才是真正人工智能爆发的时刻
Li Xiang9:01
Because strictly following best practice is actually anti-human — doing whatever you want, that is what satisfies human nature. So I think a good organisation, an excellent organisation, and an excellent person, often has to fight against human nature.
因为严格的按照最佳实践 其实是反人性的 对随心所欲 然后才是 然后满足人性的 对所以呢 我觉得一个好的组织 对然后一个卓越的组织 然后一个卓越的人 很多时候其实要跟人性做对抗
Li Xiang20:08
I think within five years there will be no general agent. There will be an agentOS, making it convenient for people in each profession to develop the agent they need on this agentOS.
我觉得5年之内没有通用agent 会有一个agentOS 方便各个专业的人 在这agentOS上 其实开发出来自己需要的agent
Li Xiang53:30
There is no way to pluck out the tenth bun directly. Everyone may feel the tenth bun is what filled them up, but every bun before it cannot be skipped.
就是没有办法直接摘第10个包子 虽然可能大家觉得第10个包子吃饱了 但前面每一个包子其实都跳不过去
Li Xiang58:32
Because the stronger the model's capability, the higher the possibility that it goes rogue — just as the stronger a person's capability, the more I need their professionalism to be strong.
因为模型能力越强 也意味着他胡来的可能性越高 就跟一个人能力越强 其实我要需要他的职业性越强
Li Xiang1:06:38
When energy between people is always present, then I think these arguments, these discussions, these fights are a more complete brain. When that energy disappears, then these political discussions, these different ideas, are actually internal friction.
当人和人之间能量始终存在的时候 然后我觉得这些争执 这些讨论这些吵架 就是一个更完善的大脑 当这些能量消失的时候 然后这些政治讨论不同的想法 其实就是内耗
Li Xiang1:51:06
So what do I think wisdom is? I think wisdom is our relationship with all things.
那我觉得什么是智慧 我觉得智慧就是我们和万物的关系
Li Xiang2:27:22
Figures
| Li Auto 2025 projected revenue | over 100 billion RMB | 1:25:51 |
| Li Auto 2024 sales volume | 500,000 vehicles | 1:26:52 |
| VLA cloud VL base model scale | 32B (32 billion parameters) | 42:23 |
| Verification cost per ten thousand kilometres | from 170,000-180,000 RMB down to just over 4,000 | 1:08:38 |
| Superalignment team size | over 100 people | 1:12:40 |
| End-to-end team size | 200 people | 1:59:08 |
| Time saved for Li Auto by DeepSeek's open source | nine months | 23:09 |
Glossary
- VLA / Vision-Language-Action model
- Li Xiang defines it as the driver large model, able to understand the physical world, drive and communicate with people like a human driver.
- MoE / Mixture of Experts
- An architecture that combines a set of expert capabilities; Li Xiang uses it as a metaphor for a CEO calling on different experts.
- AgentOS
- Li Xiang's preferred description for a general agent: each profession develops its own professional agent on top of it.
- RLHF / Reinforcement Learning from Human Feedback
- Using data such as human takeovers and driving habits as feedback, so the model aligns with humans and society.
- World model / traffic physics world simulation
- The traffic world Li Auto builds through reconstruction plus generation, used for exams, generating training data and future L4 operations.
How to listen
Founders and engineers watching AI agent deployment, autonomous driving technical routes, and organisation and talent management — especially anyone who wants to know how VLA is actually trained and whether the AgentOS idea holds up.
After 2:09, the retrospective on the tenth anniversary and the part on family and intimate relationships — more personal reflection, skippable.