You Can't Build Superintelligence Inside a Data Center: The Bottleneck Is Real-World Signal
Automating AI R&D is not the same as superintelligence: Tom Reed argues that for most tasks the data only exists once you deploy models in real markets, and neither sample efficiency nor simulation can fill the gap — the singularity is stuck on signal, not compute.
The video won't play here. Listen to the audio instead:
The argument · timestamps estimated from transcript position
The bottleneck for the singularity is signal, not compute
Tom's positive claim is not that superintelligence is impossible, but that it cannot be stacked up inside a data center. His mechanism: to be good at solving problems, you need a source of problems, and that source can only be hard interaction with the real world in serial time. He compresses this into one line — ‘signal is sovereign’ — bodies need pain, companies need profit, models need real deployment. So he judges that the arrival of superintelligence depends on deploying AI into the economy expensively, painfully and slowly, not on how much more compute gets stacked up inside a lab.
— Tom ReedJaggedness won't disappear before the singularity
The current way capabilities advance does not support the claim that ‘the model is accumulating a general g’: models only get better at what they have a chance to practice, so the capability boundary is more uneven than anyone expects. In 2016 no one would have predicted a model that beats Stanford Law graduates on the bar exam yet can't play tic-tac-toe reliably. Tom thinks jaggedness is mainly a product of the training data distribution, and there is no reason to believe automating AI R&D would end it, so he proposes Toner's Law: we should expect jaggedness to persist longer than expected, even after we have adjusted for it.
— Tom ReedPatchwork RL won't buy you general intelligence
Tom zeroes in on the current toolbox for capability progress: custom RL environments are the main engine. As Mechanize puts it, each environment is a single engineer spending a week patching one specific failure observed on a real frontier model; as Zhengdong Wang puts it, scaling used to hand you a bundle of capabilities at once, and now they have to be unlocked one by one with effort. Tom's challenge: are we really going to do the singularity one problem per week at a time? If RL genuinely generalised, AI companies should already have the evidence; when Dwarkesh kept pressing Dario, every one of Dario's rebuttals was just a weak analogy between RL and the pretraining generalisation of the GPT-2 era.
— Tom ReedFor most tasks the relevant data is zero
The common rebuttal is that LLMs are wildly sample-inefficient and could improve enormously, and AI bulls point smugly at DeepSeek reaching GPT-4 level on a fraction of the compute. Tom doesn't buy it: Anson Ho and Beren Millidge both point out that gains in intelligence per FLOP come mainly from data quality rather than algorithm quality, and pure algorithmic innovation yields fairly weak returns at constant compute. More fundamentally, for most tasks the relevant data simply doesn't exist — improving sample efficiency by orders of magnitude still multiplies zero. His test: if the internet had never existed, how much progress could today's talent and compute make on writing code? His guess is tens of billions of human-hours.
— Tom ReedText about a task is not the task
For most economically interesting tasks, what's in the corpus is description, commentary or advice, not a record of performing the task itself. Dominion is a board game with a huge body of strategy writing, yet Claude plays it badly: the training process that teaches it to write beautiful articles about Dominion is a different thing from building an agent that can win, because the optimisation target is ‘the advice looks good’ rather than ‘the strategy wins’. Tom's analogy is that the LLM's relationship to most tasks is a bit like McKinsey's relationship to running a company. The exception is writing code — the tokens are themselves a constituent part of the task — which is why training AlphaZero on chess commentary rather than chess games would be the anomaly.
— Tom ReedThe score goes up, the real capability doesn't
How does self-improvement inside a data center know it has actually gotten stronger? Of the submissions on SWE-bench accepted by AI automated graders, close to half are rejected by the real human maintainers of the relevant codebase — the score can be pumped up while the merge rate doesn't move. Tom asks in return: how is Claude 8 supposed to tell, from the results of a Become-POTUS eval alone, whether its brilliant new algorithm actually improved Claude 9's ability to become president? Evals are useful, but they are always stubbornly incomplete proxies; in the end the only eval that counts is voters, customers, the economy and the P&L. Without those, your singularity gets ruthlessly Goodharted.
— Tom ReedMarket data isn't hidden, it doesn't exist yet
The second layer of doubt is more fundamental and more speculative: even an arbitrarily intelligent simulator cannot generate practice data for many tasks, because that data can only be produced through interaction with the relevant market. This is Hayekian: that something is good is a signal generated by millions of actors revealing preferences through behaviour, and it cannot be simulated inside a data center — not because data collection happens to fail, but because it doesn't exist anywhere yet. He uses his own board gaming as an example: a game of Twilight Imperium takes 12 hours, he'll plan weeks ahead until he feels certain of winning, and about 5 minutes into the real game his predictive power goes to zero. Continual learning doesn't save you either, because without a genuinely diverse problem set you can't tell which algorithm is better.
— Tom ReedThe winner may be the deployer itself
Corollary one: a policy focus on ‘internal deployment’ and ‘automating AI R&D’ may be misplaced — superintelligence isn't impossible, it just won't arrive in the shape currently imagined, on the timeline currently predicted; the next phase looks more like the grind of deployment, customer discovery and real data collection. Corollary two: data holders have far more power than they realise (Ukraine is one example), and the asymmetry between ‘companies that can build frontier models’ and ‘entities holding proprietary deployment data from real economic activity’ may flip. Corollary three: labs may conquer vertical industries one at a time, and the way to win is to become the deployer yourself.
— Tom ReedIn their own words · checked verbatim
Put simply, I think it’s impossible to get good at solving problems without access to a source of those problems, and that arduous interaction with the real world in serial time is the only such source.
Tom Reed1:41
We should expect jaggedness to last longer than we expect, even after we adjust for jaggedness lasting longer than we expect. Call this Toner’s law.
Tom Reed5:01
As Mechanize describe it, each environment is the product of a single engineer spending a week patching a single observed failure in a frontier model.
Tom Reed6:01
I would say LLMs relate to most tasks kind of like McKinsey does to running a company.
Tom Reed10:36
This means that, for most tasks, our approach to training LLMs is akin to trying to train AlphaZero on chess commentary rather than on games of chess. My take would be that even a 100% success rate at predicting the tokens of a Robert Caro biography doesn’t generalise to becoming POTUS.
Tom Reed11:40
At the end of the day, the only eval that matters is the voter, the customer, the economy, and the bottom line. Without access to any of these, your singularity will be mercilessly Goodharted.
Tom Reed14:06
You can’t simulate this from inside a data centre, because the inputs aren’t being withheld through any contingent failure of data collection — they just don’t exist anywhere yet, and they won’t until the market generates them.
Tom Reed15:28
I don’t think superintelligence is impossible — it just won’t arrive in the shape we currently anticipate, or on the timelines currently forecasted.
Tom Reed20:13
Figures
| Share of SWE-bench submissions accepted by AI automated graders that are rejected by real human maintainers | close to half | 14:06 |
| Input required to build one RL environment | one engineer for a week, patching one specific failure of a frontier model | 6:01 |
| Human labour needed to reach current coding capability if the internet had never existed | tens of billions of hours | 9:31 |
| Length of a game of Twilight Imperium | about 12 hours | 16:41 |
| Point at which the pre-game Twilight Imperium planning model loses predictive power | about 5 minutes into the game | 16:41 |
| Timeframe in which the authors of AI 2027 expect fully automated coding to produce systems that comprehensively surpass humans | within one year | 0:00 |
Glossary
- sample efficiency
- Reaching the same capability with less data; Tom thinks it can't rescue tasks where the relevant data is zero.
- Goodhart Singularity
- Tom's term for a fake singularity produced by self-improving only against an evaluation metric.
- jaggedness
- The same model performing wildly differently across tasks, with uneven capabilities.
- continual learning
- The ability to learn on the job after deployment, seen as an escape route around the data bottleneck.
- RL environments
- Task scenarios custom-built to train a model; Tom says they are patched together by hand, one at a time.
How to listen
Investors and policy researchers who care about AGI timelines, and founders in vertical industries who hold proprietary data and want to know where they sit in the AI value chain.
5:01–7:15 covers jaggedness and RL environments; skip if you already know this material.