Anti-Data Centres Is Almost Entirely a Chinese Psyop
The American left and right are both fighting new data centres, and Thomas Sohmers thinks it is almost entirely a Chinese psyop — China is building flat out while the West slows itself down with regulation.
The video won't play here. Listen to the audio instead:
The argument · tap a timestamp to hear it
Training is compute-bound, inference is memory-bound
Thomas lays out the fundamental difference between training and inference clearly: in training you already have the whole corpus, so you can feed tokens in massively parallel — it is essentially a compute-bound problem, which is why export-control frameworks mostly target FLOPS. But inference is generative: you do not know the token five words ahead, so you have to generate one at a time in sequence, and every token generated means reading the weights again. That turns it into an extremely memory-bound problem that cannot be parallelised the way training is. This difference dictates that inference hardware has to be designed on completely different logic from training chips.
— Thomas SohmersCompute grew 120x, memory bandwidth only 17x
Over the decade from 2014 to 2024, the FLOPS of a single NVIDIA GPU rose about 120x, but memory bandwidth rose only 17x. Thomas says this is the real manifestation of the memory wall: the ratio of compute to memory bandwidth has diverged enormously, and the further you go, the worse the memory-bound problem gets. One reason is that SRAM cells have scaled far more slowly than ordinary transistors over the past fifteen years, and the architecture has barely changed in thirty or forty years. Another reason is incentive — in the CNN era of the 2010s, models were essentially compute-bound, so you just threw more FLOPS at them, and nobody had any incentive to solve memory bandwidth.
— Thomas SohmersThe margin on selling cached tokens is ‘obscene’
Thomas says the cost of processing a cached token is roughly one-thousandth of recomputing and generating that token, so the margin on selling cached tokens is absurdly high. Nearly every provider charges a higher rate for cache than for ordinary input tokens, then a lower rate on reads; users feel it is cheap, but providers are earning ‘obscene margin’ on cache reads. He notes Anthropic's API business has been reported at 80 points of gross margin, and he is not surprised — getting 80 points in any industry is hard. He thinks those margins will compress with competition, and that if OpenAI and Anthropic stopped training, they would be substantially profitable overnight.
— Thomas SohmersAnti-data-centre sentiment is almost entirely a Chinese psyop
Thomas says the left and right have now reached near-consensus against data centres, which in his eyes is the most frightening thing in politics, and he thinks it is almost entirely a Chinese psyop. He rebuts the water-consumption argument: an in-and-out burger uses more water than the largest data centre in the US, golf courses are orders of magnitude higher still, and these data centres are closed-loop liquid-cooled systems that in many cases do not want to use water at all. His argument: China itself is building flat out, adding multiple gigawatts of generation capacity, and displacing homes and forcing blackouts for training capacity, while the West slows itself down on the basis of false information.
— Thomas SohmersData centres will not take residents' electricity
Thomas says utilities have to provide an offer of energy availability, and that offer is already baked into everyone's costs and capabilities, so a data centre cannot possibly pull power from electricity already allocated to residents — it is impossible. Moreover, the data centres being built today come with generation capacity covering their own use and more, but are simply not allowed to connect to the grid — if they were connected, they would actually lower everyone's electricity price. He also points out that utilities are lobbying against new generation capacity, because market forces would let more capacity push prices down, which is good for consumers but not for utilities.
— Thomas SohmersKV cache turns quadratic cost into linear storage
Thomas explains the KV cache mechanism: the Transformer originally recomputed all previous tokens for every token generated, until it was realised you do not need to redo the parts already processed, so the K and V matrices are stored. The key point is that under attention, the compute per token grows quadratically with sequence length, while storing K and V grows only linearly. So KV cache trades storage for computation that gets expensive fast. But storage is itself complexity: every user has a unique KV cache, and you have to decide how long to keep it and how to manage it across a large system.
— Thomas Sohmers96% of agentic coding tokens are cache hits
Thomas mentions SemiAnalysis's agent X benchmark, based on a large number of cloud code sessions with dozens to hundreds of turns, plus sub-agents. They found that in these real traced code-generation, agentic coding sessions, roughly 96% of tokens are cache hits. That means if you know a workload has an extremely high cache rate, the importance of cache retrieval rises sharply, because those caches become very large. He gives an example: GPT-4 was leaked as a 1.8 trillion parameter model, about 900 GB of model weights after quantization; if Claude is 10 trillion parameters, that is about 5 TB of weights. But under long context, a single user session can reach the 100 GB scale, so 50 users' context exceeds the model weights themselves.
— Thomas SohmersChinese labs used algorithms to route around memory export controls
Thomas says context length is heading from today's million tokens toward 10x or more, which is very hard under quadratic memory cost. Algorithmic progress over the past year has been interesting, able to further reduce the storage and compute context requires, and Chinese model labs have innovated a lot here. He cites DeepSeek V3: because export controls limit chips with the highest memory capacity and FLOPS, they innovated methods that do not need those chips, using multi-head latent attention to shrink KV cache size substantially, at the cost of spending more FLOPS. He says the major US labs, as far as he knows, have not used MLA, but that will change in the future.
— Thomas SohmersIn their own words · checked verbatim
You make all of your money on selling cashed input and output tokens.
Thomas Sohmers11:16
Like if they stop training, they'd be massively profitable overnight.
Thomas Sohmers12:17
The scariest thing to me on the political spectrum and the way all of this being treated is that it's now become a almost unifying issue on left and right about being anti-data centers. I think that is almost entirely a Chinese psyop.
Thomas Sohmers21:45
Like, you know, a single in and out uses, you know, more more water than, you know, the largest data centers in the United States.
Thomas Sohmers22:49
about 96% of all the tokens that go through these entire sessions are cached.
Thomas Sohmers39:37
the value per unit of intelligence is probably closer to 1,000 fold, not just the 60 fold you're talking about.
Thomas Sohmers1:02:26
Figures
| Positron Series C raise | $875 million | 0:00 |
| Positron valuation | $5 billion | 0:00 |
| 2014-2024 FLOPS gain per NVIDIA GPU | about 120x | 7:08 |
| 2014-2024 memory bandwidth gain | 17x | 7:08 |
| Anthropic API business gross margin | 80 points | 11:16 |
| Cache hit rate in agentic coding sessions | about 96% | 38:31 |
| Leaked GPT-4 parameter count | 1.8 trillion | 38:31 |
| GPT-6 Astra accuracy on the Ruler long-context benchmark | over 95% | 59:19 |
| GPT 5.6 accuracy on the Ruler long-context benchmark | about 70% | 59:19 |
Glossary
- KV cache
- Stores the K and V matrices in attention to avoid recomputing already-processed tokens.
- memory wall
- Compute improves far faster than memory bandwidth, worsening the memory-bound problem.
- MLA
- The attention mechanism DeepSeek proposed to compress KV cache, trading more FLOPS for a smaller cache.
- quantization
- Lowering numerical precision (e.g. FP16 to FP4) to compress models and caches, potentially at some cost in accuracy.
- Ruler
- A benchmark testing a model's ability to retrieve randomly inserted information in long contexts.
How to listen
Founders and investors watching AI infrastructure, inference economics and geopolitics — especially teams building chips, cloud services or agent products.
The first 4 minutes of host intro and ads are skippable.