The world is too loud. Read what matters.

The a16z Show

AWS deliberately doesn't sell all GPUs to top AI labs; it reserves allocations for startups

AWS revenue contains 30–40% from companies that were once startups; this is Garman's explanation for why AWS deliberately reserves GPU quotas for small companies even when compute is scarce and could all be sold in a day to large AI labs.

Compute allocationCustom siliconEnterprise agent deploymentCapital expenditureSupply-chain constraintsBedrock

The video won't play here. Listen to the audio instead:

Density concentrates in the middle and later sections: GPU allocation logic, supply-chain bottleneck whack-a-mole dynamics, the origins of custom chips, and enterprise agent trust barriers. The first ten minutes lean toward historical review and can be fast-forwarded.

The argument · tap a timestamp to hear it

4:07

Today's startups are tomorrow's big customers

Early on, AWS conducted internal research concluding that the customers most worth serving were startups. Garman notes that 30–40% of AWS revenue today comes from companies that were once startups. This is the business logic behind AWS's deliberate decision to reserve GPU quotas for small companies even when compute is scarce—not charity, but because today's two-or-three-person team might become tomorrow's major bank or government customer willing to pay premium prices for performance, and also the source through which AWS learns new requirements from large enterprises.

— Matt Garman
18:17

Forgo revenue now to reserve GPU quota for startups

Garman admits AWS could sell every single GPU to top AI labs and clear inventory in a day, but the company deliberately doesn't, to keep the ecosystem healthy. He reveals that for GPU requests from customers, AWS ultimately fulfills roughly 60% in some form—sometimes by changing the region, configuration, or delivery timing. Simultaneously, AWS announced plans to acquire 2 million NVIDIA GPUs over the coming years, the largest single compute expansion announced to date.

— Matt Garman
21:24

Customer diversification is AWS's defense against AI bubble claims

When asked whether AI is a bubble, Garman's answer is concentration data: some cloud providers tie 30–60% of their capacity to one or two major customers. AWS's largest single-customer share is only single-digit percentage, usually less. His logic: even if some venture-backed startups valued at $1 billion ultimately fail, most of AWS's compute revenue comes from production workloads that have already proven positive return on investment. That demand won't disappear if a few companies collapse.

— Matt Garman
26:26

Bottlenecks shift endlessly; they're never fully resolved

Garman mentions reading *The Goal* in college, which taught him that there's never a single bottleneck, only the latest one at any given moment. Power might be the constraint this month, then next month it shifts to memory, TSMC capacity, HBM, or networking components—even a connector shortage. He also recalls Thailand's 2011 flooding that triggered a global hard-drive shortage, illustrating how AWS manages its supply chain four or five vendors deep, tracking thousands of individual parts and replenishing whatever gets tight.

— Matt Garman
32:33

Custom silicon started with virtualization tax, not AI hype

The Trainium and Graviton story didn't start with AI; it began thirteen or fourteen years ago when customers complained about the "virtualization tax" and wanted bare-metal performance. AWS first offloaded network virtualization to a separate card, then acquired the Annapurna team (which built ARM processor cards) and offloaded storage virtualization too—and ended up with Graviton as a side effect. Today Graviton costs 20% less and runs 20% faster than competing chips; over 90% of the top 100 customers use it. Trainium followed the same "prove it, then scale" path; the third generation is already in production through next year. Garman admits the name was a miss—a chip called "training" now runs more inference than training.

— Matt Garman
39:40

Most enterprise agents today still lack autonomy by design

Garman observes that most agents deployed in enterprises today are simple and non-autonomous; someone remains in the loop to sign off. He finds that enterprises commonly make the mistake of handing agents five human steps as-is, when they should instead tear it down and rebuild: let agents run many approaches in parallel and converge on the answer, solving problems the way computers do rather than copying human workflows. What really stops enterprises from letting go isn't lack of technical capability—it's trust. They fear agents with unchecked permissions will accidentally delete production databases. Garman finds this caution "probably pretty reasonable."

— Matt Garman
44:44

Bedrock's moat: customer data never leaves their VPC or reaches model makers

Garman reiterates that Bedrock has promised since its inception that customer data never leaves their VPC and model providers never see prompt content—a position criticized three years ago as "AWS moving too slowly on AI." But as customers shift from proof-of-concept to production, this conservative architecture choice has become an advantage. OpenAI and Anthropic workloads are now migrating to Bedrock because enterprises view their own data as their most valuable asset and won't hand it to model vendors.

— Matt Garman
52:59

Inside AWS: engineers manage agent teams, not code directly

Inside AWS's own "frontier teams," writing code is no longer assistant autocompletion—it's fully handed to agents. Engineers now manage an agent team instead, which has fundamentally changed how fast new products ship. Organization structure is shifting too: where it once took ten people long-term to own a product direction, now three or four people deliver the same output. Whether teams should stay smaller and switch projects more often is an experiment AWS is running internally. Garman's conclusion is measured: humans managing agents and agents managing humans will coexist long-term.

— Matt Garman

In their own words · checked verbatim

That's why agentic workflows tend to perform better on AWS than anywhere else.

Matt Garman8:11

I saw recently that we say, you know, yes, in some way, shape or form to something like 60% of the requests we eventually get.

Matt Garman18:17

we recently announced we're going to be buying, you know, 2 million NVIDIA GPUs over the next couple of years

Matt Garman19:21

We're, you know, single digit percentages at the highest and usually it's less than that.

Matt Garman21:24

it turns out there's never one constraint. There's always just the latest constraint.

Matt Garman26:26

We have a guarantee that your data never leaves your VPC.

Matt Garman44:44

It really is agent first, the agents write all of the code. You're just managing a team of agents and driving that.

Matt Garman52:59

I'll say that agents manage people, people manage agents.

Matt Garman53:59

Figures

AWS revenue from companies that began as startups30–40%4:07
AWS 2026 capital expenditure$220 billion16:15
GPU request fulfillment rateApproximately 60% of requests ultimately fulfilled in some form18:17
NVIDIA GPUs planned for acquisition in coming years2 million19:21
Graviton cost and performance advantage20% lower cost, 20% higher performance35:36
Top 100 customers using GravitonOver 90%35:36
Annual per-capita tax reduction in data center countiesApproximately $5,000 less per person annually30:32

Glossary

Firecracker
AWS's custom lightweight virtualization technology with strong security isolation and fast startup; commonly used by agent sandboxing vendors.
Bedrock
AWS's managed foundation model service that guarantees customer data never leaves their VPC and is not shared with model providers.
Graviton
AWS's custom ARM-based server processor, designed to deliver 20% lower cost and 20% higher performance than competing chips.
Trainium
AWS's custom AI acceleration chip originally designed for training but now predominantly used for inference workloads.
FDE (Frontier Deployment Engineer)
AWS technical teams deployed on-site at customer locations to teach them self-assessment frameworks over weeks, then withdraw.

How to listen

Who it's for

Founders, investors, and technical leaders focused on how hyperscale cloud providers allocate GPU capacity, custom chip strategy, and enterprise-agent deployment barriers.

Skip

The opening review of AWS's early history and startup landscape changes has low information density and can be skipped.