The world is too loud. Read what matters.

张小珺·商业访谈录

Tesla FSD Is Only Assisted Driving's GPT-3 Moment, Not L4

End-to-end let FSD jump several versions within a year, but one takeover every 45 minutes versus Waymo's one critical takeover per 100,000 km is roughly a 1000x gap. For L2 the goal is the optimal takeover rate, not the lowest one.

Autonomous DrivingTesla FSDEnd-to-EndL4Waymo
Meng Xing spans three vantage points — founder, investor, and COO of Didi autonomous driving — and lays out the two product systems of L2 and L4 thoroughly. Information density is high, and the back half's estimate of what it costs to replicate FSD is especially valuable.

The argument · tap a timestamp to hear it

11:17

Assisted driving is mid-product; driverless is still Pre-A

Meng Xing slices by product stage rather than funding round: most Chinese autonomous driving companies were founded in 2016 or 2017, have run for 8 years, and may be very late in funding rounds, yet their product stage is still Pre-A or A. Assisted driving is already mid-product; driverless has yet to produce a product with a closed upper bound, revenue potential, and large-scale rollout. The gauge is user scale — the number of people who actually ride a driverless car every day, who have ridden more than a hundred times, is far, far smaller than for ride-hailing products. So information in this industry mostly circulates among insiders, and much of what outsiders hear is secondhand; trace it back to the source and that person may never have experienced it at all.

— Meng Xing
21:18

V12 changed architecture without losing ground — that alone is remarkable

What Meng Xing drove most was 11.3.6, the last non-end-to-end version before V12 shipped, and 12.3 after it shipped. His personal judgment differs from much of what's said online: 12.3 and the last V11 version are actually not that different. But from a technical standpoint, V12 swapped in a completely new architecture and performance didn't decline — it stayed consistent or even improved slightly, which is already extremely remarkable, because people usually expect an architecture change to deliver a tenfold improvement before it counts as success. 12.3's trait is strong responsiveness: it won't pretend not to see a scenario, but its success rate isn't high — during rush hour on the highway with dense merging traffic it basically fails 100% of the time; conversely, it can also engage in many scenarios you wouldn't expect, like a backyard or a cattle pen.

— Meng Xing
29:09

For L2 the takeover rate isn't lower-is-better; there's an optimum

This is the most counterintuitive point in the whole piece. Meng Xing has discussed it with quite a few product leads and founders who work on both L2 and L4, and the view he largely agrees with is: before L2 achieves extremely high driving capability (essentially L5), there exists an optimal takeover rate — neither too high nor too low. The reason is that safety under human-machine co-driving is guaranteed by two capabilities together — the car's autonomous driving ability, and the human's ability to step in and backstop. As takeover mileage rises, the human's backstop ability actually falls: if you drive three highway trips and never take over, you certainly won't keep both hands, both feet, and your eyes on the road the whole time. So net capability doesn't necessarily rise. What the optimum is, nobody knows — some say 50 km, some say 100 km, some say 5,000 km. L4 is an entirely different system: it must absolutely be the best, with no takeovers at all, because it doesn't assume a human backstop.

— Meng Xing
33:36

The L1-to-L5 naming muddles capability and responsibility together

This naming was proposed by the US highway administration; the smaller the number, the more human involvement is required, and L5 requires no human involvement at all. But it actually speaks to two things: responsibility and capability. L4 means no human involvement at all in a limited environment — if something goes wrong it's not the human's responsibility but the system's; at the same time it implies the car has the capability to drive safely in most situations. The problem is that when people bring it up they conflate the two — when automakers advertise L2 or L3 they're talking about capability, but the definition defines responsibility, and responsibility should generally be defined by regulation. So an automaker can say it has L4 capability but ships it as an L2 product that complies with L2 rules, and if something goes wrong the responsibility lies with the driver. For both L2 and L3 the responsibility is the driver's; for L4 it's the platform's.

— Meng Xing
52:20

V12 is assisted driving's GPT-3 moment, not its ChatGPT moment

Meng Xing states plainly that V12 is not yet autonomous driving's ChatGPT moment; if you must draw an analogy, it's more like assisted driving's GPT-3 moment. A GPT-3 moment means: after scaling the model to a certain size, some capabilities begin to emerge, and they don't recede; capabilities you hadn't anticipated emerge; and you confirm this is the right path to continue on. V12 did indeed surface new capabilities — for instance suddenly crossing a line where it seemingly shouldn't but with higher efficiency, or automatically adjusting to avoid obstacles in a parking lot — none of which were preconfigured or even specifically trained, but naturally acquired from what was buried in historical data. But it's still not that product that drops all at once, breaks out of the circle, works for everyone, and gets universally good reviews.

— Meng Xing
54:11

The FSD team is just over 300 people, trading imperfection for speed

As of the end of May this year, Tesla's entire FSD team was just over 300 people — the largest moment in FSD's history, having been under 200 for quite a long stretch. Meng Xing gives two reasons: first, Musk's management style, wanting a big runway ahead, few people, little bureaucracy, so the boss's demands and execution can reach the lowest level directly, with a very flat structure; second, accepting imperfection — up to version 12.4, highway and low-speed were still two separate versions, and much half-finished engineering was simply halted because there weren't enough hands. This is the same pattern as early Facebook's Ship Fast & Break Things — allow imperfection, but move fast. Tesla certainly has the money to hire ten thousand people; it chooses not to.

— Meng Xing
1:17:32

Replicating FSD's first-generation version costs $1-3 billion

Meng Xing breaks down the elements needed to build today's FSD: team, training infrastructure, data collection. He thinks people aren't the most expensive part — the most expensive is training infrastructure: GPUs, data centers. Tesla up to 12.3 was probably operating at a scale of 10,000 H100s; 10,000 H100s corresponds to roughly $300 million of investment, and he estimates half of that, i.e. over $100 million, went into data centers. Adding data collection costs (assuming you already have that many cars out collecting data), the first-generation version is roughly $1-3 billion of investment, spread over several years. If Tesla spends 10 dollars, replicating it for 3 to 4 dollars is a normal ratio.

— Meng Xing
1:19:18

In China nearly everyone is following, but with heavier rule-based backstops

First-tier new forces, NIO, XPeng, Li Auto, Huawei, Momenta and others are nearly all following end-to-end; some already have concrete launch plans or have already shipped, and every one has adjusted its org structure and talent pipeline. But Meng Xing judges that Chinese companies' end-to-end differs from Tesla's in one respect: the core modules are similar, but Chinese companies are relatively more conservative and add some rule-based backstop modules, heavier than Tesla's, and are less willing to leave that entirely to the user to backstop. Also, China's low-speed scenarios are vastly more complex than America's; without end-to-end you might not reach a shippable product at all, so for Chinese companies this has become inevitable, and the investment may be more resolute. On generational gap, Tesla shipped this year and everyone else may be one to two years behind, but Tesla hasn't stopped, so it's hard to catch up.

— Meng Xing

In their own words · checked verbatim

There used to be a joke in our industry: people kept asking when driverless would arrive, and the answer was always five years from now, and then five years after that, and it always seemed like it would never happen. But actually by 2020 it happened — Waymo had already opened up a point in time where you could hail a ride from them, at least it was in the progressive tense. It was no longer a future-tense kind of work.

过去有一个笑话 在我们这个行业里面 一直说无人驾驶 什么时候实现 然后这个答案 永远是五年之后 五年之后又是五年之后 然后永远好像实现不了 但其实到2020年实现了 Waymo已经公开 给大家打车的一个时间点 至少是进行式了 它已经不是 未来式的一个工作

Meng Xing14:13

A takeover every 45 minutes — does that count as passing, as a product experience? For everyone there's really no universal standard. A lot of people have no way to judge. And that's true — for pure driverless, for people doing L4, a takeover every 45 minutes is simply unacceptable, it's just too bad. It's not a question of whether it's acceptable; it's completely unacceptable. But for assisted driving, this may just represent the level of the time.

45分钟出现一次接管 这到底算不算过关这个产品体验对吧 就是对于大家来讲其实没有平凡标准的 很多人其实无从去认为 那确实也是 比如对于纯无人驾驶来讲 就是对于做L4的人来讲 45分钟一次接管 这简直是不可接受的一件事情 就是太差了 不是说能不能接受的问题 这是完全不可接受的一件事 但对辅助驾驶来讲 其实这可能就是属于代表了当时的水平

Meng Xing23:19

Actually as L2, until it achieves an extremely, extremely high driving capability — extremely, extremely high basically means what we'd call L5, the same as what human driving is — before that there is actually an optimal takeover rate, and this optimal takeover rate is neither too high nor too low.

其实作为L2 直到它实现了极高极高的 车的驾驶能力之前 极高极高基本上是我们意义上的L5 这就是跟人类驾驶是什么一样 之前其实它是有一个最优接管率的 这个最优接管率不是过高也不是过低

Meng Xing30:22

But think about it — what kind of person would keep staring at this autonomous driving system, guessing whether it can solve a problem, and having to take over when it can't? In our industry that's called a safety test driver, and we have to pay that test driver. It's actually a pretty grueling job.

但是你想啊 什么样人一直要盯着这个自动驾驶系统 然后猜他能不能解决一个问题 且在解决不得的时候还要去接管 这个在我们行业里面的就叫安全测试员 这个测试员是我们要付钱给他 是一个其实蛮辛苦的工作

Meng Xing37:24

What is a GPT-3 moment? A GPT-3 moment is when I scale the model — expand it to a certain size — and it starts to show what we call emergent capabilities, and they don't recede. It emerged with capabilities I hadn't anticipated before, and I believe this path is the right one and can be continued.

GPT3时刻是什么呢 GPT3时刻是说 我把模型scale到 就是扩量到一定大之后 它开始所谓我们叫 涌现的一些能力 而且它没有后撤 它涌现了一些 我以前没有预想到的能力 且我认为这条路是正道 可以继续

Meng Xing53:32

But I think a tenfold improvement in a year — I probably won't see that. 12.3 is too far off; you haven't even started on it yet, so you don't know what 12.3 is, and there's no way to evaluate it.

但我觉得一年提升比如十倍 我觉得这个我可能看不到 12.3太远了 你都还没开始去呢 所以不知道12.3是什么 这个也无法评价

Meng Xing1:12:41

Actually we need to see evidence. You invest this much — if you can put in 100 million, roughly what kind of result can you see? If you put in 500 million, what kind of result can you see? I think this is a process I see people somewhat lacking today — they just feel this is a good direction, Tesla just bought machines and put 300 million dollars in first, and then people just hurry to catch up, we at least have to catch up, and that's how everyone goes about it.

其实我们需要看到证据的 你投入这么多 能投入一个亿 大概能看到一个什么样的结果 投入五个亿能看到一个什么样的结果 我觉得这个是我今天看到大家有点缺失的一个过程 就觉得这是一个好方向 特斯拉光买机器 把三个亿美金先投进去了 然后人反正就赶紧去 我们至少得追上 大家这么去干

Meng Xing1:15:43

Figures

FSD team size (as of end of May this year)just over 300 people54:11
Tesla training compute (before 12.3)10,000 H100s1:17:32
Investment corresponding to 10,000 H100sabout $300 million1:17:32
Estimated investment to replicate FSD's first-generation version$1-3 billion1:17:32
Cost ratio of replicating FSD relative to TeslaTesla spends 10 dollars, replication costs 3 to 4 dollars1:17:32
Waymo San Francisco indoor critical takeover mileageabout once per 100,000 km27:19
Waymo Phoenix critical takeover mileageabout once per 300,000 km27:19
FSD 12.5 general-scenario takeover rate300 to 500 miles or km27:19
Tesla FSD North America penetration ratesingle-digit percentage1:29:54
FSD subscription price (after price cut)$991:29:54

Glossary

MPI / MPCI / takeover mileage / critical takeover mileage
The mileage driven between two takeovers; MPCI counts only critical takeovers where not taking over would mean a crash.
BEV + Transformer
A perception architecture that fuses multiple sensors into a bird's-eye view and uses a Transformer for temporal tracking.
VLM / vision-language model
A multimodal large model; Li Auto uses it for System-2-style judgment, but running it on-device has about a one-second delay.
Remote Assist
When an L4 vehicle hits a scenario it can't handle on its own (like being stopped by traffic police), it requests instructions from a remote human.
garbage in garbage out
Data quality determines model quality; data cleaning and processing capability matter more than the model itself.

How to listen

Who it's for

Engineers and product leads working on autonomous or assisted driving, plus investors and founders watching Tesla FSD and the end-to-end route.

Skip

The first ~8 minutes comparing China and US startup experiences has little to do with the autonomous driving throughline.