The world is too loud. Read what matters.

老石谈芯

Computing power is becoming the costliest part of the car—this is the core wager driving Li Xiang's chip investment

Li Xiang projects per-vehicle compute costs have already exceeded $2,000 and continue climbing toward $4,000, $8,000, or $10,000; once AI computing power becomes the car's core capability, both that cost and capability must remain in the automaker's own hands.

Custom chipsDataflow architectureEmbodied intelligenceAutomotive chipsTeam management
To see how automakers truly cost out custom chip development: cost curves, architectural tradeoffs, and organizational mechanisms are explained concretely, not in slogans.

The argument · tap a timestamp to hear it

6:04

Computing power costs approach battery costs, forcing the self-research decision

Xie Yan says Li Xiang calculates that compute costs alone for a single vehicle will exceed $2,000 and keep rising—becoming $4,000, $8,000, even $10,000—potentially matching battery costs as the vehicle's largest component. More fundamentally, if AI capability is the core of an embodied-intelligence car, then the automaker must control its foundation itself—just as leading automakers in the fuel-engine era didn't fully outsource engines and chassis to suppliers. Otherwise there's no differentiation from competitors.

— Xie Yan
14:22

Publishing the chip architecture publicly is not a replication risk

Li Xiang published M100's reconfigurable dataflow architecture to the ISCA conference. Internally, some questioned why proprietary work should go public. Xie's reasoning: the gap from paper to actually building the chip, running large models efficiently, and perfecting the compiler and operators is huge—others cannot easily replicate it. Besides, technology advances regardless; keeping secrets won't prevent others from catching up. The only way is to move faster ourselves. Publishing also attracts peer feedback and brings more people together to develop this path.

— Xie Yan
18:26

Removing all caches shifts complexity entirely to the compiler

The M100 NPU has no traditional multi-level caches; all on-chip storage is SRAM managed by software logic, with no hardware automation. The benefit: eliminates the scheduling and cache-coherence logic that consumes 30–40% of CPU transistors, yielding simpler, more efficient hardware. The cost: the compiler must do what hardware normally does, becoming extremely difficult to write. Xie summarizes this as: dataflow architecture transfers the complexity von Neumann hardware bears into software and the compiler, solving problems at the static compilation stage.

— Xie Yan
29:31

Dataflow chips have no state machines, like Schrödinger's chip

Xie explains the fundamental difference between dataflow architecture and Turing-machine-style von Neumann: the latter has a definite state each cycle and can pause to inspect at any time; dataflow chips compute on data arrival, not time or cycle triggers. Slice the system at any moment and its state is uncertain—he calls this 'Schrödinger's chip.' Debugging becomes much harder than traditional architectures, yet the final computation's precision requirements remain identical.

— Xie Yan
47:42

A 45-billion-transistor chip succeeded on its first tape-out

M100 is a large 5-nanometer, 400-square-millimeter chip with 45 billion transistors, using an entirely novel architecture—it lit up on first tape-out. From tape-out return to running Linux took under 12 hours; the team refused to leave the lab, even at midnight. Xie calls this single success remarkable, not because iteration was minor, but because architecture, compiler, OS, and model teams were bound together in co-design from day one.

— Xie Yan
50:44

Model algorithms converge; competitive moats ultimately rest on hardware

Xie predicts that over time, model-algorithm ideas will converge—one person's insight is another's, and talent movement pushes everyone the same direction. Software-level moats are at best training costs, not thought leadership. True future differentiation comes from hardware, or software designed with hardware as one. This moat resembles Apple: everyone knows Apple is great, but replication is hard because Apple controls its own chip, OS, and application ecosystem. Even with the chip, you can't build iPhone.

— Xie Yan
54:47

The chip compresses end-to-end response time to 0.28 seconds

On M100, Xie says the vehicle runs Mach's native multimodal VLA model with end-to-end response down to 0.28 seconds. He compares this to human reaction time from photon to retina to action: faster reaction means safer. With fine-grained processing and small intervals, braking and steering become fluid without jank—comfort and safety tie directly to response speed.

— Xie Yan
1:10:02

Small teams enable paradoxically better collaboration than large ones

On hardware-software co-development, Xie makes a counterintuitive point: keep teams small. Large teams want to become self-contained and resist collaboration, having enough internal power. If a task theoretically needs 10 people, the optimal configuration is 8 or fewer—each person's efficiency peaks. At 15–20 people, internal sub-units emerge and collaboration fractures. He likens this resource-gathering problem-solving to lighting a campfire.

— Xie Yan

In their own words · checked verbatim

In a dataflow architecture, if you cut through it at any given moment, its state is not fixed—it's a bit like Schrödinger's chip.

其实在数据的架构里面 你切一刀的话 在某一个时刻切下去 它的状态是不固定的 有点像是薛定谔的芯片

Xie Yan1:01

A single vehicle might already face compute costs exceeding $2,000, and we believe it will keep climbing—it's not just $2,000, not even a starting point. It could become $4,000, maybe $8,000, maybe $10,000.

比如说当时可能一辆车光算力的成本 就要花掉2000美元以上 而且我们觉得是会越来越多 就是不是2000美元 不是一个只是个开始 它可能会变成4000 可能会变成8000 1万

Xie Yan6:04

How are we qualified to make a chip? And from the start, we set extremely high targets—not just to clone or replicate. Our goal was several times better performance and lower cost than buying off-the-shelf, achieving that level of leadership. Honestly, most people weren't certain we could pull it off.

我们何德何德能做芯片 而且是一开始我们就提出来非常高的指标 不是说做一个芯片 只是克隆一下复制一下 而是我们当然提的目标 就是一个好几倍的性能 比外购更低的成本 做到这种领先的程度 其实大部分人心里是不确定的

Xie Yan9:08

I've always believed that technology keeps advancing regardless of whether we keep things hidden—others will catch up anyway. So the only way is for us to run faster.

我是一直觉得技术就是在进步 并不会因为我们把一些东西藏起来 别人就不会追上来 所以唯一的方法就是我们跑得更快

Xie Yan14:22

Once you've figured something out from first principles, if it can be done, it will inevitably be done.

你从第一心上想明白的一件事情 如果它可以被做到 它就必将被做到

Xie Yan37:36

This kind of moat is a bit like Apple. Everyone knows Apple is great, but it's extremely hard to replicate Apple, because Apple has its own chip, its own operating system, its own application ecosystem. To replicate the entire package is extremely difficult.

这种壁垒有点像苹果 大家都知道苹果好 但是你要复制给苹果很难 因为苹果有自己芯片 自己操作系统 自己的应用体系 你要把整套都复制了 非常难

Xie Yan50:44

The approach is that teams should be small. Once teams grow large, they want to become self-contained and don't want to collaborate with others, because they have enough power on their own.

方法是 我觉得团队规模要小 团队规模大了之后 他就希望自成体系 然后不希望跟别人合作 因为我有足够的力量

Xie Yan1:10:02

Figures

M100 peak compute1280 TOPS15:23
M100 actual efficiency~82%48:42
NPU compute units56 fully homogeneous15:23
CPU cores24 A78 cores15:23
Memory bandwidthLPDDR5X, 273 GB/s15:23
End-to-end latency0.28 seconds54:47

Glossary

dataflow architecture
Chip architecture where computation is triggered by data arrival rather than instruction sequence execution
test time training (TTT)
A technique where models continue training or updating parameters during inference and runtime operation
just in time compiler
Dynamic compilation of code during program runtime, rather than static compilation before execution
software hardware co-design
Simultaneous design and iteration of chip hardware and software systems from inception

How to listen

Who it's for

Auto executives and chip architects, plus investors tracking embodied intelligence and self-driving hardware roadmaps.

Skip

Around the 29-minute mark, the analogy about dimensions is a bit convoluted and can be skipped without affecting the core argument.