L3 Is Not an Extension of L2 but a Precursor to L4
Li Auto's autonomous driving team used end-to-end to cut 2 million lines of code down to 200,000, and longitudinal control beat every competitor for the first time — not because the people got smarter, but because they finally built capability instead of features.
The video won't play here. Listen to the audio instead:
The argument · tap a timestamp to hear it
The HD map path is a dead end because China has 9.7 million km of ordinary roads
Early autonomous driving assumed laying virtual tracks on the ground: HD maps plus lidar all over the car, running like a high-speed train. But highways are only just over 300,000 km; the remaining 9.7 million km are ordinary roads, and it is impossible to map them all. Even if you did, ordinary roads get dug up today and rerouted tomorrow, and updates cannot keep up. So around 2018 Tesla was the first to say no HD maps, no lidar, pure vision — the direction was seen very correctly, very precisely, very early. At the time some in the industry did not understand; some understood but had already invested too much money in heavy lidar and could not turn around.
— Lang XianpengThe key to BEV is not stitching images together but fusing first, then projecting back
The naive approach is for each camera to recognize on its own and then stitch the results together — that is late fusion. The problem is you cannot guarantee every frame is recognized correctly: this side misses a person, that side catches it, whom do you trust; one car has half extracted here and half there, and when stitched together is it 1.5 meters or 3 meters. Tesla's BEV first does one unified computation over all image information, recovers the complete object in the real world, then projects it back into each image, so consistency is guaranteed. Lang Xianpeng uses the mass-energy equation as an analogy: the person who first thought matter and energy could convert was the most impressive; whether you use fission or fusion comes later.
— Lang XianpengDefining autonomous driving by scenarios cannot be exhausted and the scenarios fight each other
Many startups boast about defining all scenarios with rigorous means: weather split into clear and rain, rain split into heavy, medium and light; traffic split into congested and dense; lighting split into night and overcast; then multiply by ego speed and the other party's state, and the combinations run into tens of millions. There are three problems: first, there are too many scenarios to exhaust; second, scenarios are not fully orthogonal, so fixing one right-turn scenario may affect another right turn on a rainy day with heavy traffic; third, the long tail is unimaginable — a horse running out mid-road, the road surface suddenly collapsing into a big pothole — these cannot be defined. So the path of dividing autonomous driving by scenarios does not hold up in itself.
— Lang XianpengL3 is not an extension of L2 but a precursor to L4
Lang Xianpeng believes L2's scenario design is never complete and never fully solvable — solve one and it affects another — so L3 is not L2 plus plus plus plus; whether in scenarios, approach or essence, it is not an extension. L3 and L4 are developed as one: L4 is very sexy but you cannot announce nationwide L4 in one step, so you first have to get the full park-to-park scenario capability working, covering campuses, cities, highways, gates and ETC; when capability is insufficient a human supervises, the share of no-takeover segments grows, and finally A to B is all no-takeover — that is L4.
— Lang XianpengOne takeover per 200 km is the threshold for launching L3
Li Auto's key measured metric is MPI takeover rate: one takeover per 350 km on highways, one per 50 km in cities, and combined about one takeover per 200 km, at which point you can launch L3 supervised autonomous driving. At forty or fifty km a day, that is roughly one takeover a week, and that takeover is not necessarily a safety takeover — it may just be that the ride felt uncomfortable. The safety requirement is higher: ten times safer than human driving. This number turns ‘when can we launch L3’ from a feeling into a measurable threshold.
— Lang XianpengCutting 2 million lines of code to 200,000, longitudinal control beat every competitor for the first time
On April 15 they did a closed-door development sprint in Zhongguancun, and a little over a month later Lang Xianpeng drove from Zhongguancun to Beijing Jiaotong University and confirmed not a single line of rules was used. The earlier mapless and light-map versions were 2 million lines of code; after end-to-end only under 200,000 remained, a 90% reduction. What surprised him more was longitudinal acceleration and deceleration: before, on the mapless and light-map versions, no matter how they tuned it they could not brake to a smooth stop like a human; end-to-end on its first time in the car was better than every autonomous driving car he had ridden in before, including competitors. He asked the person next to him whether they had tuned it with rules, and the answer was they really had not — it was learned.
— Lang XianpengThe supplier demanded Li Auto disband its autonomous driving team before it would keep working with them
In 2021, while doing the upgraded Li ONE, the ADAS supplier knew Li Auto wanted to do an upgrade; not only did they demand an expensive development fee, they refused white-box delivery, and even attached the condition that Li Auto hand all future autonomous driving R&D to them and disband its own autonomous driving team. Lang Xianpeng's exact words were that the supplier felt you absolutely had to kneel down and beg them, because there was no time left. His response was that even if he died, he would die standing. On the first day of Chinese New Year, Li Xiang called to ask whether he had the resolve; he said he had come three years ago waiting for this moment, and if he could not do it he would resign to take responsibility, and Li Xiang immediately created a group chat and announced in-house development.
— Lang XianpengThe foundation model is the groundwork; without it you work for someone else
Lang Xianpeng judged that Li Auto must build its own foundation large model, with reasons split into necessity and feasibility. There are three necessities: first, supply risk — the foundation model is like the groundwork and the cornerstone, and if one day they stop giving it to you, no matter how well you build on top it is a castle in the air; second, a general foundation model adapts to all customers and will not optimize to your requirements, so you cannot get the top position in the AI field; third, once it is truly built well, it is a top company and everyone else may end up working for it. He believes AI will not bloom in a hundred flowers but will ultimately consolidate into the hands of a small number of companies that truly have foundation large model capability.
— Lang XianpengIn their own words · checked verbatim
Our earlier mapless and light-map versions were 2 million lines of code; now only under 200,000 remain — we cut 90% of the code.
我们之前无图版本和清图版本的代码 这一条行数是200万行代码 现在只剩下不到20万行 我们缩减了90%的代码
Lang Xianpeng1:28:14
I asked whether they had used rules or tuned it somehow, and he said, boss, we really did not tune it — it was learned.
我说你们是不是用规则 或者怎么调呢 他说老婆我们真没调 就是学出来的
Lang Xianpeng1:29:14
The method that seems stupid is actually the shortest path, because you have to go through the hardship or the experience once yourself before you know why it absolutely had to come to this.
就是你貌似笨的方法 其实它是最捷径的方式 因为你必须要把前面 吃的苦 或经历的东西 你要经历一遍 你才知道 它是为什么一定要到现在的
Lang Xianpeng1:39:22
Figures
| Tesla per-vehicle sensor plus chip cost | 1000 USD, about six or seven thousand RMB | 15:16 |
| Tesla in-house chip compute | 144 TOPS | 9:08 |
| Takeover-rate threshold for Li Auto's L3 launch | Combined one takeover per 200 km (350 on highways, 50 in cities) | 57:51 |
| Change in lines of code after end-to-end | From 2 million lines down to under 200,000, a 90% reduction | 1:28:14 |
| Li Auto autonomous driving team size | Over 800 people | 1:43:23 |
| Tesla autonomous driving team size | 200-plus to 300 people | 1:43:23 |
| Training scale of the first end-to-end model | Under 1 million clips; at the time it was 600,000 or 800,000, he cannot remember exactly | 1:25:12 |
| Li Auto annual autonomous driving investment | Several billion RMB | 1:37:21 |
Glossary
- BEV
- Bird Eye View, a perception representation that unifies multiple camera feeds into a god's-eye view.
- Transformer
- A general network architecture proposed by Google that can vectorize and encode inputs such as images.
- VLM
- A bolt-on model that can recognize and understand scene elements, handling scenarios end-to-end has not seen.
- MPI
- Mean kilometers per intervention — how many kilometers on average before a human must take over, a key metric for autonomous driving capability.
- VLA
- Vision-Language-Action model, the form in which autonomous driving capability grows out of a foundation model; Lang Xianpeng says it will be seen within one to three years.
- 3DGS
- 3D Gaussian Splatting, reconstructing real scenes with Gaussian points, used to generate simulated test questions for exams.
How to listen
Engineers working on autonomous driving or embodied intelligence, investors watching the intelligent driving supply chain, and founders who want to know what end-to-end actually changed.
The first 10 minutes of technical explainer are basic; if you already know BEV and the lidar comparison, fast-forward to the 50-minute mark.