The world is too loud. Read what matters.

SemiAnalysis Weekly

Nvidia's Flagship GPU Memory Regresses for the First Time: Why 4-HI HBM Wins

Rubin Ultra's HBM was cut from a promised 1TB to 192GB — less than the 288GB shipping today. This isn't a technical step backward, it's capacity rationing under the DRAM shortage: 4-HI stacks yield better and can squeeze out more than twice as many HBM cubes.

ComputeHBMNvidiaMemorySupply Chain

The video won't play here. Listen to the audio instead:

An episode that thoroughly works through the HBM height debate: from yields and bandwidth-per-capacity pricing to why model parameter counts stopped exploding. High information density, and the second half's reasoning is worth more than the opening.

The argument · tap a timestamp to hear it

2:04

Rubin Ultra was cut from 1TB to 192GB

At GTC last year, when Nvidia previewed Rubin Ultra, it said 1TB of HBM per package: 4 compute dies, 16 HBM4E stacks, 32 gigabit (4GB) per die, 16-HI stacks. A year and a half later, the actual configuration is 192GB of HBM4: only 2 compute dies, 3GB per die of HBM4, 8-HI stacks. That is lower than the 288GB of Blackwell Ultra and vanilla Rubin — the first time in Nvidia's flagship GPU history that a next generation has less capacity than the one before it.

— Myron Xie
4:11

The downgrade was forced by DRAM supply

The motive is mainly supply. DRAM wafer capacity has barely grown, while accelerator shipments are rising and HBM demand is rising with them. Suppliers and customers like Nvidia, Google and Broadcom ran the numbers: the logic capacity locked in at TSMC is fixed, and if that HBM is made into 12-HI cubes, there aren't enough cubes to pair with all the logic. Switch to 8-HI and the same wafers yield more cubes, and the equation balances. This is capacity rationing, not a performance trade-off.

— Myron Xie
5:13

Past a capacity threshold, returns diminish

Inference genuinely needs enough HBM capacity to hold model weights and the KV cache for effective batching. But once you cross that threshold, extra capacity has diminishing returns. At the same time, HBM is getting more expensive from next year because of the shortage, so buying capacity you don't use is punished harder. The assumption that more capacity is always better does not hold on the cost curve.

— Myron Xie
9:13

Parameter counts haven't kept up with system scale

In the Hopper era, an 8-GPU H200 node had 640GB of HBM in total, and Llama 3.1 405B quantized took about 60% of it. Today's largest, Kimi K3, is 2.8 trillion parameters, roughly six times Llama 3.1, but MXFP4 quantization halves the capacity requirement. Put on an NVL72 (72 GPUs × 288GB), one set of K3 weights takes only about 8% of HBM capacity. System scale is growing faster than models, so capacity pressure has actually gotten smaller.

— Myron Xie
13:39

Post-training and inference eat most of the compute

Parameter counts haven't exploded because performance can come from elsewhere: more reasoning, more post-training and reinforcement learning. And post-training looks a lot like inference, likewise favoring bandwidth over capacity. The result is that inference-sensitive compute (post-training plus standard inference) is rapidly taking over, while classic pre-training — the part that genuinely needs large capacity — is a shrinking share of total compute at frontier labs.

— Myron Xie
16:46

4-HI is a free lunch on bandwidth

HBM bandwidth is set by the interface between the cube and the SoC. HBM4/4E has 2048 data lines, and each cube supports at most 512, so you need at least 4 DRAM layers to run at full bandwidth. But suppliers charge by capacity (dollars per GB), so 8-HI and 12-HI pay two to three times as much for the same bandwidth. 4-HI has the best dollars-per-bandwidth, and as long as capacity suffices, it is a near-free lunch. Physically you also cannot go below 4-HI.

— Myron Xie
21:56

4-HI yields better and doubles the cubes

Every layer added to an HBM stack carries a yield loss. Assume 99% yield per layer: a 1% loss compounded 8 or 12 times is far higher than compounding it only 4 times. 4-HI also has a longer production history and a simpler manufacturing flow, and it is easier to deliver power into the stack. So the number of harvestable HBM cubes from 4-HI is more than double that of 8-HI, and the bottleneck shifts from HBM to logic wafers, substrates and PCBs.

— Myron Xie
32:08

The memory shortage won't ease this decade

Why aren't memory makers, who are making a fortune, expanding capacity? Semiconductor manufacturing has extremely long lead times: you first build the cleanroom (the number one reason short-term capacity can't be added), then fill it with equipment, and the equipment itself is constrained by ASML's EUV and related tools. ASML can only make so many a year, because its own supply chain — for instance the supplier that makes extremely smooth mirrors — is also long. Myron's TLDR: the memory shortage will not ease within this decade.

— Myron Xie

In their own words · checked verbatim

the presumption from there was that you keep going taller

Myron Xie2:04

once you get beyond that threshold what we find is that the returns to that are diminishing

Myron Xie5:13

the dollar per bandwidth proposition is much better off for 4 high

Myron Xie18:50

if it's like 99% yield for each, each layer, this is a simple example, the 99, that 1% yield loss compounded 8 times or 12 times, it's much higher than the yield loss if you compound that only 4 times

Myron Xie22:56

in terms of maximizing tokens per HBM wafers, this is like going to fall high is how you sort of best optimize for that

Myron Xie28:00

the loudest cries for fall high are coming from the labs

Myron Xie34:14

Figures

Kimi K3 parameter count2.8 trillion9:13
Share of HBM capacity taken by one set of K3 weights on NVL72about 8%10:18
Data lines per cube for HBM4/4E204816:46
Time for the memory shortage to easenot within this decade (before 2030)32:08

Glossary

HBM / High Bandwidth Memory
Stacked DRAM that provides high-bandwidth memory for AI accelerators.
4-HI / 8-HI / 12-HI
The number of DRAM dies stacked in an HBM cube; more layers means more capacity but lower yield.
KV cache
Caches attention keys and values during inference, consuming HBM capacity and enabling batching.
MXFP4
A 4-bit floating-point quantization format that halves the capacity requirement for model weights.
scale-up domain
A cluster of GPUs joined by high-speed interconnects into a single memory pool, such as NVL72.

How to listen

Who it's for

Engineers, investors and founders watching the AI compute supply chain, especially anyone trying to understand how HBM supply constraints are reshaping GPU specs and the memory makers' landscape.

Skip

The small talk and article background at the start, 0:00-1:01, can be skipped.