Nvidia's Flagship GPU Memory Regresses for the First Time: Why 4-HI HBM Wins
Rubin Ultra's HBM was cut from a promised 1TB to 192GB — less than the 288GB shipping today. This isn't a technical step backward, it's capacity rationing under the DRAM shortage: 4-HI stacks yield better and can squeeze out more than twice as many HBM cubes.
The video won't play here. Listen to the audio instead:
The argument · tap a timestamp to hear it
Rubin Ultra was cut from 1TB to 192GB
At GTC last year, when Nvidia previewed Rubin Ultra, it said 1TB of HBM per package: 4 compute dies, 16 HBM4E stacks, 32 gigabit (4GB) per die, 16-HI stacks. A year and a half later, the actual configuration is 192GB of HBM4: only 2 compute dies, 3GB per die of HBM4, 8-HI stacks. That is lower than the 288GB of Blackwell Ultra and vanilla Rubin — the first time in Nvidia's flagship GPU history that a next generation has less capacity than the one before it.
— Myron XieThe downgrade was forced by DRAM supply
The motive is mainly supply. DRAM wafer capacity has barely grown, while accelerator shipments are rising and HBM demand is rising with them. Suppliers and customers like Nvidia, Google and Broadcom ran the numbers: the logic capacity locked in at TSMC is fixed, and if that HBM is made into 12-HI cubes, there aren't enough cubes to pair with all the logic. Switch to 8-HI and the same wafers yield more cubes, and the equation balances. This is capacity rationing, not a performance trade-off.
— Myron XiePast a capacity threshold, returns diminish
Inference genuinely needs enough HBM capacity to hold model weights and the KV cache for effective batching. But once you cross that threshold, extra capacity has diminishing returns. At the same time, HBM is getting more expensive from next year because of the shortage, so buying capacity you don't use is punished harder. The assumption that more capacity is always better does not hold on the cost curve.
— Myron XieParameter counts haven't kept up with system scale
In the Hopper era, an 8-GPU H200 node had 640GB of HBM in total, and Llama 3.1 405B quantized took about 60% of it. Today's largest, Kimi K3, is 2.8 trillion parameters, roughly six times Llama 3.1, but MXFP4 quantization halves the capacity requirement. Put on an NVL72 (72 GPUs × 288GB), one set of K3 weights takes only about 8% of HBM capacity. System scale is growing faster than models, so capacity pressure has actually gotten smaller.
— Myron XiePost-training and inference eat most of the compute
Parameter counts haven't exploded because performance can come from elsewhere: more reasoning, more post-training and reinforcement learning. And post-training looks a lot like inference, likewise favoring bandwidth over capacity. The result is that inference-sensitive compute (post-training plus standard inference) is rapidly taking over, while classic pre-training — the part that genuinely needs large capacity — is a shrinking share of total compute at frontier labs.
— Myron Xie4-HI is a free lunch on bandwidth
HBM bandwidth is set by the interface between the cube and the SoC. HBM4/4E has 2048 data lines, and each cube supports at most 512, so you need at least 4 DRAM layers to run at full bandwidth. But suppliers charge by capacity (dollars per GB), so 8-HI and 12-HI pay two to three times as much for the same bandwidth. 4-HI has the best dollars-per-bandwidth, and as long as capacity suffices, it is a near-free lunch. Physically you also cannot go below 4-HI.
— Myron Xie4-HI yields better and doubles the cubes
Every layer added to an HBM stack carries a yield loss. Assume 99% yield per layer: a 1% loss compounded 8 or 12 times is far higher than compounding it only 4 times. 4-HI also has a longer production history and a simpler manufacturing flow, and it is easier to deliver power into the stack. So the number of harvestable HBM cubes from 4-HI is more than double that of 8-HI, and the bottleneck shifts from HBM to logic wafers, substrates and PCBs.
— Myron XieThe memory shortage won't ease this decade
Why aren't memory makers, who are making a fortune, expanding capacity? Semiconductor manufacturing has extremely long lead times: you first build the cleanroom (the number one reason short-term capacity can't be added), then fill it with equipment, and the equipment itself is constrained by ASML's EUV and related tools. ASML can only make so many a year, because its own supply chain — for instance the supplier that makes extremely smooth mirrors — is also long. Myron's TLDR: the memory shortage will not ease within this decade.
— Myron XieIn their own words · checked verbatim
the presumption from there was that you keep going taller
Myron Xie2:04
once you get beyond that threshold what we find is that the returns to that are diminishing
Myron Xie5:13
the dollar per bandwidth proposition is much better off for 4 high
Myron Xie18:50
if it's like 99% yield for each, each layer, this is a simple example, the 99, that 1% yield loss compounded 8 times or 12 times, it's much higher than the yield loss if you compound that only 4 times
Myron Xie22:56
in terms of maximizing tokens per HBM wafers, this is like going to fall high is how you sort of best optimize for that
Myron Xie28:00
the loudest cries for fall high are coming from the labs
Myron Xie34:14
Figures
| Kimi K3 parameter count | 2.8 trillion | 9:13 |
| Share of HBM capacity taken by one set of K3 weights on NVL72 | about 8% | 10:18 |
| Data lines per cube for HBM4/4E | 2048 | 16:46 |
| Time for the memory shortage to ease | not within this decade (before 2030) | 32:08 |
Glossary
- HBM / High Bandwidth Memory
- Stacked DRAM that provides high-bandwidth memory for AI accelerators.
- 4-HI / 8-HI / 12-HI
- The number of DRAM dies stacked in an HBM cube; more layers means more capacity but lower yield.
- KV cache
- Caches attention keys and values during inference, consuming HBM capacity and enabling batching.
- MXFP4
- A 4-bit floating-point quantization format that halves the capacity requirement for model weights.
- scale-up domain
- A cluster of GPUs joined by high-speed interconnects into a single memory pool, such as NVL72.
How to listen
Engineers, investors and founders watching the AI compute supply chain, especially anyone trying to understand how HBM supply constraints are reshaping GPU specs and the memory makers' landscape.
The small talk and article background at the start, 0:00-1:01, can be skipped.