AI's bottleneck isn't the model — it's copper and transformers
The model is no longer the bottleneck. Everything "south of the model" is: chips, memory, power, cooling, all the way down to copper mines. Demand is infinite while supply is pinned by physics and regulation — and for the first time, capital converts directly into compute.
The video won't play here. Listen to the audio instead:
The argument · tap a timestamp to hear it
The bottleneck has moved south of the model
Ben says this is the most important technology in history, so it needs an entirely new infrastructure stack: not just new chips and new system software, but new ways of delivering power, and even replacing copper. Ragu adds the key judgment: over the past three years model capability rose steadily, so the model is no longer the bottleneck — the bottleneck is "everything south of the model." Martin backs it with founder flow: hardware projects used to be maybe 5% or even 3% of what top founders did, now it's over 20% to 30%. Founders as a group are far smarter than VCs as a group, and they've already decided this is where the innovation is.
— Ben Horowitz / Ragul Raguram / Martin CasadoDemand is infinite, so you only worry about margin
For judging whether demand really exceeds supply, Ragu gives two signals: first, the people who know demand best are placing enormous purchase orders — hyperscaler capex is about $700 billion this year and reportedly heading to $1 trillion combined next year; second, they can see demand coming from frontier labs, AI data companies, enterprises, and the US and internationally. Martin adds a counterintuitive mechanism: in a normal business you worry about growth, but in AI, because demand is infinite and growth is infinite, what you worry about is gross margin — that is, efficiency. And many efficiency problems are really physical limits of hardware: existing systems weren't built for AI workloads.
— Ragul Raguram / Martin CasadoChip prices are actually going up
Martin says this is a signal he's never seen: chip prices have always gone down, the price curve goes like this and then back like that, and now prices are rising. On the supply side, essentially everything is booked through 2028, and there have even been multi-day auctions for a few thousand GPUs. He also gives the double expansion on the demand side: first, tokens consumed per unit of work are exploding, with chat or agents going from 100 tokens to several thousand; second, the number of users will expand from developers to all knowledge workers and beyond. Compare the internet era, when most of what got laid down was speculative dark fiber — today every GPU built is already pre-sold.
— Martin CasadoThe memory in your servers can pay for the cloud migration
Martin tells a CFO anecdote: a large public company had historically been very resistant to moving to the cloud and had accumulated a lot of servers. When it did an inventory count, it found memory in those servers had appreciated so much that selling it would cover the entire cost of migrating to the cloud. Ragu follows with a harder number: at Hot Chips, a leading memory vendor said that the demand in hand today alone would take three years of capacity to satisfy — and that's just today's demand, not future demand. Martin's summary: power, cooling, memory, GPUs — name it, we're short of it.
— Martin Casado / Ragul RaguramFor the first time, money converts directly into compute
Martin says building things used to be an engineering problem: you could throw engineers at it, but you were bound by engineering physics — that's where The Mythical Man-Month comes from. Now it's more like a resource constraint: pour money into the system, the system produces results, and the bottleneck is the system's ability to match the resources you put in. Ben pushes this to the organizational level: the startup world used to know that if I'm two years ahead and you hire a thousand engineers to catch up, you'll just wreck your company — nine people can't make a baby in one month. But now that play works: not hiring a hundred thousand people, but spending $3 billion to light up a massive cluster, and Grok or Kimi can suddenly appear. Everyone is still psychologically adjusting to this.
— Martin Casado / Ben HorowitzBuilding an ASIC for one model is starting to pay
Martin offers a mental model: training a frontier model today costs $3 to $5 billion, and inference has to return at least that much — assume 2x, so $10 billion. If you can save 20% on inference efficiency, that's $2 billion, and $2 billion is enough to build an ASIC. So the industry has reached the point for the first time where building an ASIC for a single model pays off. And unlike traditional software, model weights are fixed — there isn't that much state and dynamism. He isn't sure the world moves to per-model ASICs, but the model explains why architectures are increasingly customized for these enormous capital commitments.
— Martin CasadoRack power pushes AC out of the picture
Rack power is going from roughly 5 to 10 kilowatts toward 100 to 250 kilowatts, compute density is up about 70x, and cooling has to go from air to liquid. Martin says at this power level AC is no longer usable and you must go to DC — and DC itself needs cooling and is extremely dangerous, which is ironic given Edison sold DC back then by demonstrating how dangerous AC was. Then there's the political environment: liquid cooling isn't enough, it has to be environmentally friendly liquid cooling; having power isn't enough, you have to contribute power to society rather than draw it away. About 10% of data centers waste enormous amounts of water and act as parasites on the grid, and that will end.
— Martin CasadoOnly 2% of electricians understand DC
Martin says the way data centers are built is thoroughly obsolete: with racks this dense, floor load ratings and soundproof walls all have to be redone, otherwise you disturb the neighbors and no state will allow it. He also flags an overlooked cost item — besides memory prices, one of the fastest-rising areas is reinforced concrete. The harder constraint is people: 800-volt voltages in data centers are extremely dangerous, and only 2% of American electrical engineers or electricians are certified for DC. Meta has free training programs to fill the gap, but that also shows AI will destroy jobs while creating a lot of new electricians.
— Martin CasadoThe 2028 power gap doesn't add up
By 2028, new data centers will need about 44 gigawatts of additional power, while the grid is expected to add only about 25 gigawatts. Martin explains how big a gigawatt is: Flagstaff, where he grew up, has forty or fifty thousand people and uses less than a gigawatt, so one gigawatt can basically light and air-condition an entire town. Why can't utilities and hyperscalers build faster? You need people to build, you need permits, you need power you can actually plug into, and you have to build your own plants — and transformers and turbines are in short supply. This is not a software problem, not something you solve by having engineers work the weekend.
— Martin CasadoBanning data centers in the US exports the opportunity
Martin says things are bad enough that new companies often have to go to Mexico or Australia to get GPUs, because it's too hard in the US. That amounts to banning data centers and handing jobs and long-term economic opportunity to other countries. His standard: data centers should give back to the community — better power, no noise, no water problems, and added jobs — and then everyone should be held to that standard. He cites examples that already do it: some data centers self-power, send power to the state during the day, and borrow power back at night when the state doesn't need it, because plants always run at peak capacity while data centers are steady around the clock and cities peak in the day and trough at night — a symbiotic relationship.
— Martin CasadoIt shouldn't be called AI, it should be machine intelligence
Martin explains why the fund is called Machine Age: first, AI is the wrong name, it should be machine intelligence. Because this is not how humans think — it's a cache of human thinking, a collection of human thought; to this day we don't know how to make an AI with no knowledge rebuild language from scratch. Second, the formal term AI has been used in computer science for 70 years and carries science fiction and Bostrom-style baggage. Third, you have to emphasize the machine half: the people who said software eats the world eventually arrive at a place where you pour money in and get stuck on the real machines underneath.
— Martin CasadoNvidia won't go pick up silver bricks
Asked whether new companies can still cut in, Ragu says that to get 10x improvements in metrics like tokens per second per dollar, tokens per watt, tokens per rack, you need new innovation, and new innovation traditionally comes from founders reasoning from first principles and solving the problem differently. Martin adds the market logic: the incumbent chip giant is worth trillions, so even taking 5% would be a huge private company; it's not that Nvidia can't do it, it's that Nvidia is looking at the 90% that contributes the same growth. He cites Alex Rampell's story: Dan Rose said you can pick up a lot of silver bricks, but I have more gold bricks than I can pick up, so I won't even look at the silver ones.
— Ragul Raguram / Martin CasadoThese founders have to be systems founders
Ragu says this kind of founder differs from a software founder: you can't just be a researcher or a great computer scientist, you have to be able to architect and design a chip or a system, and also think through how it gets manufactured, who supplies it, and the whole downstream chain. The best founders are like Jensen, who thought through the entire ecosystem before designing the chip. Martin adds two environmental changes: first, labs are so short of supply they'll sign orders before the hardware exists, so there's an early signal; second, capital supply is loose, with plenty of money in follow-on rounds. He also says these companies' first rounds are in the hundreds of millions, with a lot of money going in before there's a product.
— Ragul Raguram / Martin CasadoFewer young founders because of the supply chain
On whether there are fewer young founders, Martin says that if what you're building has a complex supply chain, requires manufacturing and is technically complex, experience helps. Elon and Travis built software companies when they were young, and even the best founders need to accumulate experience building companies and building technology before they graduate to more complex areas with more moving parts. On one side there's someone like Michael, very young and very smart, but doing pure software AI; on the other, people like Elon or Travis with enough experience. He also says this field was neglected by the whole industry and academia for 20 years, it wasn't a growth area, so experienced people were scarce to begin with.
— Martin CasadoIn their own words · checked verbatim
Here, it goes all the way down to the copper mine. That's how widespread this thing is going to be.
Martin Casado3:04
Like we've never seen prices go like up on chips. If the prices went down, they always go down.
Martin Casado6:06
the leading memory runner said the demand they have today will take them three years of capacity
Martin Casado9:14
This, there's nothing between the money going in and then the hardware, you know, creating intelligence.
Martin Casado16:22
Nine women can't have a baby in a month. Like that's it. Like that never works. Okay. Now that works.
Ben Horowitz17:32
Only 2% of electrical engineers or electricians in the U.S. have been certified on DC power.
Martin Casado31:22
Artificial intelligence was the wrong word. Like we shouldn't have called it, it's machine intelligence.
Martin Casado38:35
It sounds like you can collect a lot of silver bricks. But I'm like, I have so many gold bricks, I can't even pick them all up.
Martin Casado41:42
Figures
| Hyperscaler capex this year | about $700 billion | 5:04 |
| Supply bookings | essentially all booked through 2028 | 6:06 |
| Expected annual token demand growth | close to 1000% | 12:16 |
| Number of developers | about 30 million | 17:32 |
| Cost to train a frontier model | $3 to $5 billion | 26:49 |
| Increase in compute density | about 70x | 28:52 |
| 2028 data center power gap | about 44 gigawatts needed, grid expected to add about 25 gigawatts | 33:31 |
Glossary
- hyperscaler
- A cloud giant that builds its own hyperscale data centers, such as AWS, Microsoft and Google.
- ASIC
- A chip customized for a specific task; more power-efficient than a general-purpose GPU but does only one thing.
- per model ASIC
- A chip designed for one specific model with fixed weights, rather than a general-purpose accelerator.
- dark fiber
- Redundant fiber laid during the dot-com bubble that was never lit up or used.
- Mythical Man Month
- The classic software engineering claim: adding people to a late project only makes it later.
How to listen
Founders and investors watching AI infrastructure, the compute supply chain, data centers and energy, plus technical leaders who want to understand why hardware startups are hot again.
The fund team introductions and disclaimers after 52:00 at the end can be skipped.