The world is too loud. Read what matters.

Asianometry

The higher HBM stacks, the slower it gets — flash stacking is coming for its spot

Every layer of capacity added to HBM dilutes bandwidth per unit; Sandisk stacks NAND the same way to build HBF, with 8-16x the capacity of HBM4, but two orders of magnitude slower reads and writes, and it has to relive Optane's software problem.

AI inferenceHBMMemory chipsSandiskMemory hierarchy
This episode lays out prefill/decode/KV Cache clearly, then follows HBF's specs, timeline and the Optane lesson downward — high information density, suited to anyone trying to understand the AI memory hierarchy.

The argument · tap a timestamp to hear it

1:03

Prefill eats compute, decode eats memory

In the prefill stage the model processes the input in parallel — user prompt, system prompt, past replies, tool calls, images, audio, internal chain of thought. Attention produces Query, Key and Value vectors for each token, and the latter two get written into the KV Cache. When prefill ends you have the KV Cache for the input, which is the precondition for computing the first output token. The decode stage then queries the KV Cache and emits tokens one at a time, inserting each new token's key and value back into the cache. The two are entirely different jobs: prefill is compute-bound because it can be parallelised; decode goes one at a time and is memory-bound, with the GPU's logic circuits left waiting for data to load from memory.

3:07

Model weights and KV Cache are exploding at once

The memory industry has to handle two exponential demands at the same time. One is model weights: Kimi K3 has 2.8 trillion total parameters, about 2.8 TB at 8-bit precision, where K2.7 was only 1 TB and DeepSeek V3 is 671 GB. The other is the KV Cache: a long-running agent runs for hours, its chain of thought is enormous, and the KV Cache grows extremely large with it. The fuller the context window gets, the more confused the agent becomes. The current context window is about 1 million tokens, equivalent to 1500 to 3000 pages of text — enough for reading documents or a large codebase, but only 11-12 hours if it is audio, and only an hour if it is HD video. Compressing context blurs the details, like compressing a JPG; another route is chunking the KV Cache and dropping rarely read chunks to SSD, but the frequently read chunks have to stay in HBM.

5:14

The higher HBM stacks, the thinner the bandwidth per unit

HBM is a 3D stack with a logic base die and 4/8/12/16 core memory dies stacked on top, with TSVs running through the central region of the dies to carry data down. The problem: the central region can only hold so many TSVs, so adding capacity necessarily dilutes bandwidth per GB, and the price of incremental capacity and bandwidth rises over time. In the memory session at Hot Chips 2026, a SemiAnalysis analyst asked directly in the Q&A after the SK hynix talk: on-chip bandwidth at the cell level is about 20 TB per square centimetre, while HBM4 is 20 layers thick, so each layer gets roughly only 20% of single-chip bandwidth — ‘you're diluting the throughput too much’. Pat Gelsinger, in front of SK hynix's vice president, called HBM ‘lousy memory’; that vice president, Kim Ho-sik, partly agreed, saying HBM is not the final answer to the Memory Wall problem but is still the best thing on the market today.

7:15

HBF stacks NAND the way HBM is stacked

High Bandwidth Flash, as the name says, stacks and connects flash dies the same way HBM stacks and connects DRAM dies, using advanced packaging to stack flash dies inside a single package — this is not the same thing as monolithic 3D NAND, which stacks memory cells. NAND was originally a hard-drive replacement, optimised for capacity; HBF raises capacity per stack by 8-16x, 512 GB against roughly 36 GB for a single HBM4 stack, while keeping HBM-class high bandwidth, and in theory it could go higher because more NAND cells are exposed to the GPU. Current HBF specs reach 3 TB/s of effective bandwidth, above Micron HBM3E's 1.2 TB/s and comparable to Micron's and Samsung's HBM4 figures. But normalised by capacity, HBF's bandwidth number is much smaller.

9:25

NAND is not DRAM: slow, wears out, and burns power

There are three costs. First, NAND reads and writes (formally, programs) are about two orders of magnitude slower than HBM, and writes are slower than reads, with latency that is fundamental; Sandisk can optimise the die to cut latency, but likely at the cost of data retention or cost per bit. Second, endurance: NAND cells wear out after a few thousand writes, while DRAM cells have essentially no endurance limit, so the number of writes into HBF has to be controlled. Third, power and heat: driving these big HBF cubes takes more electricity, which is not a small matter today; more power means more heat, which causes the system to throttle or makes the endurance problem worse.

10:25

Sandisk ropes in rivals and Google to set the standard

The idea of stacking and connecting NAND with TSVs is old: at the 2019 IEEE 3D Systems Integration Conference, Koji Sakui and Takayuki Ohba of Honda Research Institute Japan proposed stacking flash dies with bumpless TSVs, calling it High Bandwidth NAND. Sandisk formally launched HBF at its Future FWD investor day in February 2025, saying the architecture had been in development for a year with major AI players involved; the first related patent application was filed in May 2024. In July 2025 it set up a technical advisory board, joined by David Patterson and Raja Koduri. In August Sandisk announced a memorandum of understanding with SK hynix to standardise the spec together — a spec only defines how the memory interfaces with the rest of the chip, not how the stack works internally, so competitors can cooperate to grow the ecosystem, compete on their own internal product specs, and give customers a credible second source. Google and Tenstorrent also joined the HBF alliance. In August 2026, the Open Compute Project published the first public spec, version 0.7.0. Samsung has its own version, called zNAND-O.

13:28

HBF wants into Tier 1, side by side with HBM

In the memory hierarchy, Tier 0 is the GPU/TPU's SRAM, fastest but with no capacity; Tier 1 is GPU HBM, off-die but closest to the chip; Tier 2 is system DRAM, higher capacity but further from the GPU. Sandisk's idea is for GPU makers to put HBF and HBM side by side in Tier 1: HBM holds ‘hot’ data, HBF holds ‘warm’ data by virtue of capacity. The GPU interposer's ‘beachfront’ can only hold so many memory chips, and today they are all HBM; architects have to decide how much to swap for HBF. A GPU with eight stacks of 24GB HBM has about 192 GB total; swap in six HBF stacks plus two HBM stacks and the system has 3.12 TB. Because of flash endurance, what goes there should be read-heavy, write-light data — most people think model weights, though some propose the KV Cache.

15:34

Cheap memory capacity is not cheap tokens

At Hot Chips 2026, two architects, A. Agrawal and R. Giduthuri, discussed the challenges of integrating HBF into AI compute, and the core point was: systems are judged by their ability to serve tokens, and HBF's cheap memory capacity does not map one-to-one onto cheaper tokens — if data cannot be pulled out of HBF fast enough, the accelerator waits and token output suffers. So HBF is not a no-brainer; it needs a lot of software support. A paper titled “HBF Sucks?” by five researchers at Peking University plus one at Fudan points out that HBF is not a plug-and-play replacement for SSD, and software must carefully choose which data goes into HBF, such as frequently used data like shared prompts. Model type matters too: Agrawal and Giduthuri note that mixture-of-experts models may suit HBF better, because they need the capacity to hold all experts but each token reads only a few of them, so lower bandwidth per GB is acceptable.

16:36

The Optane lesson: good tech, everything else collapsed

Intel Optane (the underlying technology is 3D XPoint) was phase-change memory, non-volatile like NAND, with read/write speeds close to DRAM and a price in between; Intel positioned it as something that could work alongside CPU and flash, pitching ‘warm data’ — databases that want DRAM-class latency but are too big to fit in DRAM. The use cases were interesting, but they required application developers to write software support. Then 3D NAND rose and pushed down flash's cost per bit, gutting Optane's economics, and it never found its place between DRAM and NAND; Micron pulled out, leaving Intel alone, and nobody wanted to bet both CPU and memory on a single supplier, so it ended. This is especially relevant to HBF: the executive now leading HBF at Sandisk, EVP and CTO Alper Ilkbahar, spent five years at Intel as vice president and general manager of the Optane group. With DRAM in shortage today, some muse that Intel should have kept Optane a few more years, using its capacity and latency to hold inference model weights. But one person who dealt with Optane from the software side says that behind the scenes Optane could not do what Intel claimed. Some say HBF's advantage is that it starts from familiar flash, so it will fare better in the market; the counterargument is that Optane's PCM cells themselves worked fine and the core technology was great — what failed was the controller and go-to-market strategy, the surrounding parts. HBF has to avoid the same pit.

In their own words · checked verbatim

When you look at the internal bandwidth of a single chip, you see about 20 terabytes per square centimeter at the cell level ... You talk about going to 20 levels thick at HBM4 so maybe you are at 4 terabytes per square centimeter in a stack of 20, so you are talking about having 20% of the bandwidth of one chip [for each] layer of 20. You've diluted the throughput enormously. Why is that the correct way to go? Why are you so focused on going taller rather than going faster?

The way HBM works right now, you cannot get the additional gigabytes of capacity without reducing bandwidth per gigabyte because there are only so many TSVs available in the central area to move data.

While Sandisk can optimize the dies for lower latency - likely sacrificing either data retention or cost-per-bit in the process - this gap is fundamental. You will need to wait longer for your data than with DRAM, end of story.

A core point was that systems are judged on how well they can serve tokens ... and just because HBF has cheap memory capacity does not one-to-one mean cheaper tokens. Because if we can't get all that data out of the HBF fast enough, and that makes the accelerator wait, then token output suffers.

But someone who once worked with Optane from the software side told me that - behind the scenes - Optane just couldn't do what Intel claimed it can.

But the counterpoint to that is that Optane's actual PCM cells worked brilliantly. The core technology was great. It was most everything else - like the controller and go-to market strategy - that failed to live up to its promise.

Figures

Kimi K3 total parameter count2.8 trillion3:07
Current context windowabout 1 million tokens, equivalent to 1500-3000 pages of text4:11

Glossary

Prefill
The first stage of inference, processing the entire input in parallel and generating the KV Cache; compute-bound.
Decode
The second stage of inference, querying the KV Cache to generate tokens one at a time; memory-bandwidth-bound.
KV Cache
The cache of Key/Value vectors Attention produces for each token, avoiding recomputation.
HBF
NAND dies connected using HBM's stacking and packaging approach; large capacity but reads and writes two orders of magnitude slower.
3D XPoint / Optane
Non-volatile memory from Intel and Micron, close to DRAM in speed, which ultimately failed on economics.
mixture-of-experts
Only some experts are activated per token; needs capacity to hold all experts but reads only a few.

How to listen

Who it's for

Engineers watching AI inference cost and memory architecture, investors in chips and storage, and anyone trying to understand prefill/decode/KV Cache and the HBM bottleneck.

Skip

The opening chit-chat about Baader-Meinhof and Hot Chips sightings can be skipped.