The world is too loud. Read what matters.

硅谷101

Bypassing CUDA for inference chips is inevitable, but SRAM's window may not stay open forever

Training is a compute game, inference is a bandwidth game. Groq and Cerebras both bet on SRAM, but one discounts for determinism and the other discounts for yield; OpenAI instead chose the most expensive HBM4, because what it lacks is power, not money.

Inference chipsSRAMGroqCerebrasHeterogeneous computing
Two frontline chip entrepreneurs break down the trade-offs of the three paths taken by Groq, Cerebras and OpenAI, with plenty of engineering detail and cost accounting — for anyone trying to understand what inference chips are really betting on.

The argument · tap a timestamp to hear it

4:17

Training bets on compute, inference bets on bandwidth

In the training era no one could make money from training itself, so costs couldn't be quantified — everyone just wanted to bet on the strongest cluster. In the inference era the math is clear: how much the user pays per million tokens, and how much it costs me to produce those million tokens. More fundamentally, it's about compute density: training has massive data to feed the compute, while the decode phase of inference is strictly autoregressive — it must produce one token at a time, and each token requires reading the entire model from storage into the compute unit. So the core tension of inference shifts from compute to memory bandwidth.

— Ziyang
16:39

SRAM is the consensus, but the two companies part ways

Groq and Cerebras both saw the bandwidth demand before Transformers appeared, and both converged on SRAM — because SRAM's absolute bandwidth and bandwidth per dollar are two to three orders of magnitude higher than HBM. The cost is capacity: DRAM stores one bit with one transistor plus one capacitor, while SRAM needs six transistors, so per-chip capacity can be two orders of magnitude lower than HBM. So you have to scale up the system — hundreds or thousands of chips to hold all the model weights — and the bottleneck shifts from capacity to cluster communication and scheduling. Here the two diverge: Groq moves all scheduling ahead to compile time, while Cerebras simply stuffs a whole rack of chips into a single wafer.

— Mark
19:34

Groq's static scheduling collides with MoE's dynamism

Groq designs hardware around the compiler, pushing determinism to the extreme with no runtime scheduling overhead. But after MoE appeared, each token activates 8 out of 128 experts, and which ones are chosen is decided at runtime — unpredictable at compile time. The static approach has only two options: conservatively send all tokens and let some chips idle, or change the way the model is partitioned, but current models are hard to partition cleanly along dimensions other than experts. Idling only affects cost-effectiveness, not the fact that it's still fast, so Groq remains a good alternative to GPUs.

— Mark
25:18

The price Cerebras pays for yield

The maximum single-die size is about 800 square millimeters; a wafer can yield thirty or forty dies, and with 50% yield you still get 20 good ones. But Cerebras's entire wafer is one product — if defects exceed the 30% it can tolerate, the whole wafer is scrapped. This isn't mask cost (masks can be amortized), but the fact that out of every 10 made, maybe 9 are unusable. So it may achieve 10x Nvidia's bandwidth, but at 3x the cost of HBM, so bandwidth cost-effectiveness is discounted. Groq discounts for determinism, Cerebras discounts for yield cost, and the discounts ultimately land on token cost.

— Mark
33:58

In the inference era, CUDA will definitely be bypassed

TPU and Trainium have zero intention of being CUDA-compatible, yet people still find TPU software hard to use; Anthropic can buy in bulk because many of its engineers come from Google and know JAX. The key change: in the training era people were willing to pay for good software, but in the inference era the balance tips toward performance and cost-effectiveness. As long as it's not many hard problems but one hard problem — writing all Transformer operators on a new chip — the business logic works; plus AI coding accelerates operator optimization, greatly lowering the software stack barrier. DeepSeek would rather write low-level PTX assembly for optimization, which is exactly this logic.

— Ziyang
44:25

The faster inference is, the smarter the model gets

After agents exploded, a lot of tokens aren't for humans to read but for other agents to read; human reading speed tops out at 100 tokens per second. But the significance of inference speed isn't faster interaction — it's that in the same minute of waiting, the model can use 10x more tokens for more internal thinking, thus becoming smarter. Jensen's slide labels the horizontal axis as Smarter AI and the vertical axis as tokens per watt. Conversely, some applications that give answers quickly may just be thinking less. Time is the fundamental constraint — ask AI to predict tomorrow's stock market, and if it thinks for two days the answer is meaningless.

— Ziyang
49:04

Grasp the invariants: FFN became stable after turning into MoE

Model structure will definitely affect chip design — for example, diffusion models are compute-intensive and suit another paradigm, so you first make a heterogeneous choice: whether to do this at all. More critically, find the invariants — in Transformers the attention mechanism is still constantly innovating, with every new lab publishing papers with new ideas; but the FFN behind it has become a relative invariant since turning into MoE. RL in post-training also doesn't fundamentally affect model structure. So prioritize solving the part that has already converged, and the chip won't be left behind by algorithm iterations.

— Mark
1:09:40

Everything revolves around locality

Bill Dally's exact words are that computer architecture is like real estate: everything is about location, location, location. Temporal locality means data used once will likely be used again; spatial locality means data used once will soon be followed by its neighbors. CPU caches are applications of these two. But in the decode phase of AI inference, algorithmic temporal and spatial locality are both flattened, so you must restore locality with hardware in the architecture — putting model weights into SRAM is essentially buying the best locality.

— Mark
1:17:30

OpenAI chose HBM4 because it lacks power, not money

Jalapeño is a general-purpose inference chip that packs prefill, decode, attention and FFN all into one chip, can even run game demos, and the next generation will support training. It stacks HBM4 bandwidth to 15.4 TB per second per package, taking the most expensive path. The reason is that US data centers lack not capital and space but power, so it cares intensely about how many tokens can be produced per watt, doing a lot of work on power consumption, with tokens per megawatt greater than Nvidia's corresponding Rubin architecture. In China electricity is about 1/3 of cost, in the US about 2/3; different constraints lead to different convergence points.

— Ziyang
1:27:03

Heterogeneity inside the chip, or at the system level

OpenAI proposed the concept of Dark silicon: when a chip does different things, you can turn off or reduce power to some parts, buying power efficiency with more cost. This differs from the system-level heterogeneity of Groq and Cerebras — the latter is one chip working with other chips, putting the attention KV cache on another chip. SRAM has poor capacity cost-effectiveness, suitable for storing fixed model weights, not for storing a continuously growing KV cache. So for scenarios where power is the first constraint, OpenAI's approach is great; for scenarios demanding extremely high bandwidth cost-effectiveness, it will still converge to SRAM.

— Ziyang

In their own words · checked verbatim

Back in the training era, no one could rely on training to get any commercial return — training itself was something with no direct return.

就在训练的时代 大家是没有办法依靠训练 来获得任何商业上的回报的 就训练本身是一个 没有直接回报的事情

Ziyang5:06

This medium isn't really about the medium itself, but about focusing on the locality problem — you need to provide a higher level of locality than HBM.

这种介质其实不在于介质本身 而在于你要关注局部性的问题 你要提供 比起HBM更高一层的局部性

Mark11:11

When you can make something faster and cheaper than others, you won't tell them that your cost is actually lower.

当你可以做出一个比别人更快 而且更便宜的东西的时候 你不会告诉别人 你的成本其实是更便宜的

Ziyang22:18

In a mature company, pushing a research idea disruptively to the product team is still extremely difficult.

在一个成熟的公司里面 把一个research idea(研究想法) 颠覆式地推给产品的团队 还是难度非常大的

Mark33:29

So faster inference doesn't mean you go from a minute of thinking to interacting in a second; rather, with the same one-minute interaction time, the model can use 10x more tokens for more internal self-thinking to raise its intelligence.

所以推理速度更快并不意味着 你要从一分钟的思考 最后你跟它的交互速度 要变成一秒钟 而是在都是一分钟的 交互速度的情况下 你这个模型可以用10倍 更多的Token 做更多内部自我的思考 来提升它的智能水平

Ziyang45:36

Computer architecture is like real estate: everything is about location, location, location.

计算机系统结构就像房地产 一切都在于location location location

Figures

Single-user inference speed comparison between Groq and CerebrasGroq about 100 tokens per second, Cerebras up to 750 or even 2000 tokens per second21:18
SRAM's bandwidth cost-effectiveness advantage over HBMTwo to three orders of magnitude22:18
Maximum single-die sizeAbout 800 square millimeters25:20
Acceptable wafer defect ratio for Cerebras30%26:20
Cerebras bandwidth and cost relative to NvidiaAbout 10x bandwidth, about 3x HBM cost29:25
Amount Nvidia acqui-hired Groq for20 billion USD32:29
Target improvement of an ideal inference system over GPUsTwo to three orders of magnitude in speed, about 10x in cost-effectiveness36:32
Electricity as share of TCO for US and China data centersAbout 1/3 in China, about 2/3 in the US42:35

Glossary

arithmetic intensity
The amount of computation per unit of data read, determining whether a chip is compute-bound or bandwidth-bound.
roofline model
An architectural analysis framework that uses arithmetic intensity to judge whether an application is stuck on compute or bandwidth.
MFU
The ratio of actual compute to theoretical peak, mainly measuring utilization of the matrix multiplication portion.
MBU
The proportion of memory bandwidth actually used, a key metric for high-speed inference scenarios.
Dark silicon
Parts of a chip are turned off or powered down during specific tasks to gain power efficiency.
Amdahl's Law
Overall system speedup is limited by the unoptimized parts, explaining why heterogeneous systems only improve 10x.

How to listen

Who it's for

Engineers and founders working on AI chips, compilers or inference infrastructure, and investors trying to understand the inference cost equation and heterogeneous approaches.

Skip

From 1:13:03 to 1:14:47, Bill's startup advice to students is mostly personal reminiscence; you can fast-forward.