Bypassing CUDA for inference chips is inevitable, but SRAM's window may not stay open forever
Training is a compute game, inference is a bandwidth game. Groq and Cerebras both bet on SRAM, but one discounts for determinism and the other discounts for yield; OpenAI instead chose the most expensive HBM4, because what it lacks is power, not money.
The video won't play here. Listen to the audio instead:
The argument · tap a timestamp to hear it
Training bets on compute, inference bets on bandwidth
In the training era no one could make money from training itself, so costs couldn't be quantified — everyone just wanted to bet on the strongest cluster. In the inference era the math is clear: how much the user pays per million tokens, and how much it costs me to produce those million tokens. More fundamentally, it's about compute density: training has massive data to feed the compute, while the decode phase of inference is strictly autoregressive — it must produce one token at a time, and each token requires reading the entire model from storage into the compute unit. So the core tension of inference shifts from compute to memory bandwidth.
— ZiyangSRAM is the consensus, but the two companies part ways
Groq and Cerebras both saw the bandwidth demand before Transformers appeared, and both converged on SRAM — because SRAM's absolute bandwidth and bandwidth per dollar are two to three orders of magnitude higher than HBM. The cost is capacity: DRAM stores one bit with one transistor plus one capacitor, while SRAM needs six transistors, so per-chip capacity can be two orders of magnitude lower than HBM. So you have to scale up the system — hundreds or thousands of chips to hold all the model weights — and the bottleneck shifts from capacity to cluster communication and scheduling. Here the two diverge: Groq moves all scheduling ahead to compile time, while Cerebras simply stuffs a whole rack of chips into a single wafer.
— MarkGroq's static scheduling collides with MoE's dynamism
Groq designs hardware around the compiler, pushing determinism to the extreme with no runtime scheduling overhead. But after MoE appeared, each token activates 8 out of 128 experts, and which ones are chosen is decided at runtime — unpredictable at compile time. The static approach has only two options: conservatively send all tokens and let some chips idle, or change the way the model is partitioned, but current models are hard to partition cleanly along dimensions other than experts. Idling only affects cost-effectiveness, not the fact that it's still fast, so Groq remains a good alternative to GPUs.
— MarkThe price Cerebras pays for yield
The maximum single-die size is about 800 square millimeters; a wafer can yield thirty or forty dies, and with 50% yield you still get 20 good ones. But Cerebras's entire wafer is one product — if defects exceed the 30% it can tolerate, the whole wafer is scrapped. This isn't mask cost (masks can be amortized), but the fact that out of every 10 made, maybe 9 are unusable. So it may achieve 10x Nvidia's bandwidth, but at 3x the cost of HBM, so bandwidth cost-effectiveness is discounted. Groq discounts for determinism, Cerebras discounts for yield cost, and the discounts ultimately land on token cost.
— MarkIn the inference era, CUDA will definitely be bypassed
TPU and Trainium have zero intention of being CUDA-compatible, yet people still find TPU software hard to use; Anthropic can buy in bulk because many of its engineers come from Google and know JAX. The key change: in the training era people were willing to pay for good software, but in the inference era the balance tips toward performance and cost-effectiveness. As long as it's not many hard problems but one hard problem — writing all Transformer operators on a new chip — the business logic works; plus AI coding accelerates operator optimization, greatly lowering the software stack barrier. DeepSeek would rather write low-level PTX assembly for optimization, which is exactly this logic.
— ZiyangThe faster inference is, the smarter the model gets
After agents exploded, a lot of tokens aren't for humans to read but for other agents to read; human reading speed tops out at 100 tokens per second. But the significance of inference speed isn't faster interaction — it's that in the same minute of waiting, the model can use 10x more tokens for more internal thinking, thus becoming smarter. Jensen's slide labels the horizontal axis as Smarter AI and the vertical axis as tokens per watt. Conversely, some applications that give answers quickly may just be thinking less. Time is the fundamental constraint — ask AI to predict tomorrow's stock market, and if it thinks for two days the answer is meaningless.
— ZiyangGrasp the invariants: FFN became stable after turning into MoE
Model structure will definitely affect chip design — for example, diffusion models are compute-intensive and suit another paradigm, so you first make a heterogeneous choice: whether to do this at all. More critically, find the invariants — in Transformers the attention mechanism is still constantly innovating, with every new lab publishing papers with new ideas; but the FFN behind it has become a relative invariant since turning into MoE. RL in post-training also doesn't fundamentally affect model structure. So prioritize solving the part that has already converged, and the chip won't be left behind by algorithm iterations.
— MarkEverything revolves around locality
Bill Dally's exact words are that computer architecture is like real estate: everything is about location, location, location. Temporal locality means data used once will likely be used again; spatial locality means data used once will soon be followed by its neighbors. CPU caches are applications of these two. But in the decode phase of AI inference, algorithmic temporal and spatial locality are both flattened, so you must restore locality with hardware in the architecture — putting model weights into SRAM is essentially buying the best locality.
— MarkOpenAI chose HBM4 because it lacks power, not money
Jalapeño is a general-purpose inference chip that packs prefill, decode, attention and FFN all into one chip, can even run game demos, and the next generation will support training. It stacks HBM4 bandwidth to 15.4 TB per second per package, taking the most expensive path. The reason is that US data centers lack not capital and space but power, so it cares intensely about how many tokens can be produced per watt, doing a lot of work on power consumption, with tokens per megawatt greater than Nvidia's corresponding Rubin architecture. In China electricity is about 1/3 of cost, in the US about 2/3; different constraints lead to different convergence points.
— ZiyangHeterogeneity inside the chip, or at the system level
OpenAI proposed the concept of Dark silicon: when a chip does different things, you can turn off or reduce power to some parts, buying power efficiency with more cost. This differs from the system-level heterogeneity of Groq and Cerebras — the latter is one chip working with other chips, putting the attention KV cache on another chip. SRAM has poor capacity cost-effectiveness, suitable for storing fixed model weights, not for storing a continuously growing KV cache. So for scenarios where power is the first constraint, OpenAI's approach is great; for scenarios demanding extremely high bandwidth cost-effectiveness, it will still converge to SRAM.
— ZiyangIn their own words · checked verbatim
Back in the training era, no one could rely on training to get any commercial return — training itself was something with no direct return.
就在训练的时代 大家是没有办法依靠训练 来获得任何商业上的回报的 就训练本身是一个 没有直接回报的事情
Ziyang5:06
This medium isn't really about the medium itself, but about focusing on the locality problem — you need to provide a higher level of locality than HBM.
这种介质其实不在于介质本身 而在于你要关注局部性的问题 你要提供 比起HBM更高一层的局部性
Mark11:11
When you can make something faster and cheaper than others, you won't tell them that your cost is actually lower.
当你可以做出一个比别人更快 而且更便宜的东西的时候 你不会告诉别人 你的成本其实是更便宜的
Ziyang22:18
In a mature company, pushing a research idea disruptively to the product team is still extremely difficult.
在一个成熟的公司里面 把一个research idea(研究想法) 颠覆式地推给产品的团队 还是难度非常大的
Mark33:29
So faster inference doesn't mean you go from a minute of thinking to interacting in a second; rather, with the same one-minute interaction time, the model can use 10x more tokens for more internal self-thinking to raise its intelligence.
所以推理速度更快并不意味着 你要从一分钟的思考 最后你跟它的交互速度 要变成一秒钟 而是在都是一分钟的 交互速度的情况下 你这个模型可以用10倍 更多的Token 做更多内部自我的思考 来提升它的智能水平
Ziyang45:36
Computer architecture is like real estate: everything is about location, location, location.
计算机系统结构就像房地产 一切都在于location location location
Mark1:10:56
Figures
| Single-user inference speed comparison between Groq and Cerebras | Groq about 100 tokens per second, Cerebras up to 750 or even 2000 tokens per second | 21:18 |
| SRAM's bandwidth cost-effectiveness advantage over HBM | Two to three orders of magnitude | 22:18 |
| Maximum single-die size | About 800 square millimeters | 25:20 |
| Acceptable wafer defect ratio for Cerebras | 30% | 26:20 |
| Cerebras bandwidth and cost relative to Nvidia | About 10x bandwidth, about 3x HBM cost | 29:25 |
| Amount Nvidia acqui-hired Groq for | 20 billion USD | 32:29 |
| Target improvement of an ideal inference system over GPUs | Two to three orders of magnitude in speed, about 10x in cost-effectiveness | 36:32 |
| Electricity as share of TCO for US and China data centers | About 1/3 in China, about 2/3 in the US | 42:35 |
Glossary
- arithmetic intensity
- The amount of computation per unit of data read, determining whether a chip is compute-bound or bandwidth-bound.
- roofline model
- An architectural analysis framework that uses arithmetic intensity to judge whether an application is stuck on compute or bandwidth.
- MFU
- The ratio of actual compute to theoretical peak, mainly measuring utilization of the matrix multiplication portion.
- MBU
- The proportion of memory bandwidth actually used, a key metric for high-speed inference scenarios.
- Dark silicon
- Parts of a chip are turned off or powered down during specific tasks to gain power efficiency.
- Amdahl's Law
- Overall system speedup is limited by the unoptimized parts, explaining why heterogeneous systems only improve 10x.
How to listen
Engineers and founders working on AI chips, compilers or inference infrastructure, and investors trying to understand the inference cost equation and heterogeneous approaches.
From 1:13:03 to 1:14:47, Bill's startup advice to students is mostly personal reminiscence; you can fast-forward.