AI will hit an energy wall within three years, and the fix is to build a computer that isn't a computer
Google alone burns 3.2 quadrillion tokens a month; at 10 joules per token that's 12 gigawatts, against roughly 40 gigawatts of total US data center power. Naveen Rao thinks the only way to close that gap is to replace the underlying paradigm of the computer.
The video won't play here. Listen to the audio instead:
The argument · tap a timestamp to hear it
Energy isn't a cost line, it's the ticket to entry
Naveen's math: Google crosses 3.2 quadrillion tokens a month, and at 10 joules per token — a low estimate — that's 12 gigawatts. The US today puts about 40 gigawatts into data centers; the world, under 100 gigawatts. So one company's AI service eats more than a tenth of global data center energy. His call is that we hit the energy ceiling in about three years. More important, the ranking of scarce resources in a data center has changed: it used to be floor space, racks, networking gear, GPUs; today it's energy — get the power contract first, then figure out how to fill it. And energy is 50% of the cost of serving a token; only the rest is hardware capex and site.
— Naveen RaoWhat a squirrel does on 8 milliwatts takes a GPU watts
A human brain is about 20 watts. Scale linearly by neuron count down to a monkey brain and you get one watt, exactly your phone in your pocket. Further down, rats and bats are in the milliwatt range. A squirrel leaping between branches lands it a thousand times out of a thousand on an 8-milliwatt brain. So your phone's battery could run a hundred-plus squirrel brains. Naveen's conclusion is that biology has already built a physical substrate suited to intelligence, while synthetic systems went another way: the human cortex moves only about 16 billion bits per second, while a high-end GPU moves close to 30 trillion bits per second in and out of memory, and inside the chip it's another 10 to 100 times higher. The energy goes almost entirely into moving information, not computing on it.
— Naveen RaoMoore's Law stopped, and the abstraction layers still charge rent
Transistor counts are still rising, but frequency stopped scaling long ago, single-thread performance stopped, and now even efficiency no longer improves as process nodes shrink — Naveen says flatly that Moore's Law is basically over. His diagnosis is the abstraction layers: digital itself is an abstraction of the physical world, and transistors actually have intermediate states that get engineered into 0s and 1s. Every layer of abstraction is lossy and ignores the complexity beneath it. So their approach is to wire the physics of the semiconductor directly to the neural network, skipping the layers in between. His intuition pump is flocks of birds and ant colonies: individuals do something extremely simple, and complex behavior emerges from the whole — the field is called dynamical systems theory.
— Naveen RaoMetronomes on a board sync up on their own
Put several metronomes on a rigid board that can roll back and forth, and they synchronize to exactly the same phase, simply because each one gives the board a tiny push. That's a purely physical dynamical system. Their question was whether such a system can compute. The answer is yes — they open-sourced an image generation model called UNO built on a set of such oscillators, and it trains and produces images. Define the system state as the phases of all the oscillators, condition on the output, and generating planes, cars or birds follows different state space trajectories. It's the first demonstration that this class of system can be trained at scale and produce useful results.
— Naveen RaoThrowing away connections makes the system easier to train
Full connectivity is n squared: 10 elements and 100 connections is fine, 1,000 elements is a million connections and simply doesn't scale. Their sparsity is deliberately discarding some connections and then seeing whether the overall behavior can be rescued. The result isn't just that you can throw them away — after throwing them away the system behaves better and is more trainable, and that holds both in simulation and in real physical systems. Naveen calls this the rare case where more efficient, more scalable and better performing all hold at once, a holy grail nobody had cracked for a long time, provided you frame the problem correctly first.
— Naveen RaoThe first physical dynamics computer has already generated images
This is Naveen's first public disclosure: they built the first physical dynamics computer. The company only formally started in January of this year, with no team at the time; on June 1 they sent the design to tape-out, and the chip is back in the lab with results — the first images generated by this kind of computer. On efficiency, generating an image takes about 500 nanojoules, versus millijoules for an ordinary GPU; a nanojoule is 10 to the minus 9, several orders of magnitude apart. The reason is that it doesn't move information at all. Architecturally it differs from both CPUs and GPUs: those are von Neumann architectures with memory and compute separated and data shuttled back and forth; in this machine compute and storage are the same thing, there is no memory interface, and every compute unit is itself the memory.
— Naveen Rao4D computing: three spatial dimensions plus one of time
They call the whole thing 4D computing: physically, die stacking uses all three dimensions vertically and in plane, and the fourth dimension is time in the dynamics. The objective function is intelligence per watt. There is a thermodynamic limit that can never be crossed; animal brains are only one to two orders of magnitude from it, while today's computing systems are about 10 billion times away. Their plan is to hit the limit of 2D lithography within three and a half years, and the company's overall goal is to beat biology. Naveen expects the result to be a change in form factor: from gigawatt-scale mega data centers to small data centers scattered everywhere, more local, more adaptive, and supporting billions of robots that dynamically combine to solve problems.
— Naveen RaoA thousand times cheaper means a thousand times more consumption
Naveen closes with the Jevons paradox: when the underlying cost of an asset falls, consumption grows by more than the price drop. Halve the price and consumption more than doubles; cut the price to a thousandth and consumption grows more than a thousandfold. So he doesn't think efficiency gains shrink the market — they create the largest market in human history. Chamath then presses on the ecosystem — tape-out, packaging and so on need a whole set of people to line up. Naveen's answer is a complete product within two years, in the form of an entire rack, sold as a new data center product: tokens in, tokens out over the network, but internally structured completely differently from existing computers.
— Naveen RaoWhat ports over is the model layer, not the operator layer
Chamath asks whether existing model families and architectures can move over, given the whole industry is built around mechanisms like the KV cache. Naveen says there's a sliding scale: how much better something is determines how much migration pain people will tolerate, and his strategy is to make the payoff large enough. The port happens at the model layer, not the operator layer; existing models work, but the migration itself takes considerable compute. He points out in particular that something like matmul doesn't exist on this machine — you can describe it as matrix multiplication, but it isn't implemented as matrix multiplication, it's implemented as behavior that changes over time, where each time step can be analyzed as the current state matrix times a transition matrix. On team composition, dynamical systems theory has existed for a hundred years, so they hire from that community and they also hire people who actually build chips, and those two groups don't talk to each other — coordination is the hardest part of this company.
— Naveen RaoIn their own words · checked verbatim
So, you can imagine if models get bigger that that energy goes up and if demand grows, which it is, that energy goes up. So, we're going to run out of energy pretty fast in like 3 years or so is my estimate.
Naveen Rao5:04
Their brain runs on 8 mill of energy. You could run over a 100 squirrel brains on your phone and it it has very precise and accurate behavior.
Naveen Rao7:04
Moore's law, if you may have heard of this, it's like making transistors smaller has largely ended.
Naveen Rao10:06
This is actually the first physical dynamical computer ever built.
Naveen Rao15:09
So it's many orders of magnitude more efficient than a standard computer and it's because it just doesn't move information around.
Naveen Rao16:09
So, if you make something half the price, you'll consume more than 2x. If you cons you make something 1,000th the price, you'll consume more than 1/ 1,000th of it.
Naveen Rao19:12
Figures
| Power of Google's AI service at 10 joules per token | 12 gigawatts | 5:04 |
| Total US data center power | about 40 gigawatts | 5:04 |
| Total global data center power | under 100 gigawatts | 5:04 |
| Energy as a share of the cost of serving one token | about 50% | 6:04 |
| Human brain power | about 20 watts | 7:04 |
| Squirrel brain power | 8 milliwatts | 7:04 |
| Information transfer rate of the human cortex | about 16 billion bits per second | 8:05 |
| Bits per second in and out of memory for a high-end GPU | close to 30 trillion | 8:05 |
| Energy for a dynamics computer to generate one image | about 500 nanojoules | 16:09 |
Glossary
- dynamical systems theory
- A mathematical framework, roughly a century old, for how simple individual rules give rise to complex collective behavior.
- von Neumann architecture
- A computer structure in which memory and compute are separate and data is shuttled between them; CPUs and GPUs are both of this type.
- sparsity
- Deliberately discarding some connections in a system, which instead makes the overall behavior better, more trainable and more scalable.
- Jevons paradox
- When the efficiency of using a resource improves, total consumption rises rather than falls.
- 4D computing
- Naveen's concept: use all three spatial dimensions physically, with the fourth dimension being time in the dynamics.
- state space trajectory
- Defining the state as the phases of all the system's oscillators and watching the path it traces over time.
How to listen
Founders, investors and hardware engineers watching AI compute costs, data center power constraints and new chip architectures.
07:04 to 10:06, the biology and computer history analogies — low information density, skip to 12:07.