The world is too loud. Read what matters.

No Priors

Custom chips are a dangerous bet—when architecture shifts, rivals overtake you in nine months

Once a rival discovers a breakthrough that only works on their own chip, labs betting on custom chips could be eliminated within months—which is precisely why frontier players won't stake their future on a single proprietary chip.

Inference chipsMemory bandwidthChip architectureSemiconductor startupsAI infrastructureMarket dynamics

The video won't play here. Listen to the audio instead:

The founder speaks with rare clarity on memory bandwidth, SRAM versus DRAM tradeoffs, and chip-market game theory. Information density is high; the whole episode rewards listening.

The argument · tap a timestamp to hear it

1:01

Betting on inference speed, not training scale

Fractile was founded in summer 2022, when the industry was still being "educated" about what inference meant. Walter's bet was on the deployment era, not the training era: models would become superhuman, like AlphaGo, by investing more compute at test time. This judgment was much earlier than peers—he mentions a recently IPO'd inference chip company that was still building training chips 18 months ago, suggesting that "betting on the right direction" is itself a rare capability, not something obvious in hindsight.

— Walter Goodwin
3:02

Chip market diversity is an illusion

Walter points out that today's AI chip market looks like an extremely rich "zoo" of options, but custom chips like Google TPU, Meta MTIA, Microsoft Maya, and OpenAI Jalapeno all ultimately rely on ASIC houses like Broadcom—a $2 trillion company—to go from architecture to tape-out. They all use the same HBM memory, the same tensor cores for matrix multiplication, the same TSMC advanced packaging. Teams that truly own the full stack from architecture through physical design are rare—it's a structural feature of the industry.

— Walter Goodwin
11:24

Context length growth exposed SRAM's fundamental limits

For the first two years, the company—like Grok and Cerebrus—bet on SRAM chips: compute and weights/KV cache on the same die, extremely high bandwidth, able to push language models to thousands of tokens per second. But from late 2023 into 2024, Walter began worrying about this roadmap's scalability. AI is growing not just in parameter count, but in context length, which is becoming increasingly important. SRAM capacity is inherently limited. So the team shifted to co-developing a super-high-bandwidth DRAM solution with memory vendors and foundries. This platform ramps to production in the second half of next year.

— Walter Goodwin
12:25

Faster chatbots are just faster horses

Walter uses Henry Ford's aphorism to challenge the shallow narrative that "faster inference = smoother chat." Ford asked people what they wanted and heard "faster horses," not cars. By analogy, "faster chatbots" is just the "faster horse" of fast inference. Speed's real lever, he argues, is running trillion-parameter models at thousands of tokens per second, which radically speeds up long-running agents. That's what fast inference chips should actually target—not conversation latency.

— Walter Goodwin
22:12

The moat is a three-to-six-month architectural lead

Walter compares chip competition to frontier model races: no one can permanently own a chip that perfectly matches current workloads—in six months, needs shift again. Real advantage comes from maintaining a "deployment-ready" technology roadmap in reserve. A structural three-to-six-month lead on architecture decisions is enough to capture all new deployments in that window. That's why Fractile must own the full stack from architecture design to tape-out—not depend on external partners' timelines.

— Walter Goodwin
28:28

Bandwidth growth has lagged compute growth for two decades

Fractile internally explores "bandwidth's scaling law": in 20 years, compute improved roughly a million-fold, while memory bandwidth improved only 40-fold. This imbalance leaves room. If chips have enough bandwidth, MoE models can run much sparser—say, from 1/16 sparsity to 1/128 or 1/256—to achieve the same intelligence with less compute. Existing HBM chips hit severe bandwidth bottlenecks at ultra-high sparsity, with poor utilization. Fractile's 25x bandwidth advantage is really about pushing model architecture toward sparsity.

— Walter Goodwin
31:35

Custom chips mainly exist to negotiate NVIDIA prices down

Industry joke: the main use of big tech's custom chips is negotiating down NVIDIA prices, not doing what NVIDIA can't do—the custom chips are architecturally similar to mainstream solutions anyway. Walter says nearly all large-scale deployments now work to integrate multiple chip platforms. One reason: compute has become too critical to survival; they must maintain supply diversity rather than bet on a single technology roadmap surviving.

— Walter Goodwin
33:36

All-in on custom chips is a dangerous bet

Frontier labs won't go all-in on custom chips, Walter says, because of asymmetric payoff. If one lab bets on proprietary silicon and a rival discovers a new breakthrough that only works on their chosen chip—unlocking multiples of compute efficiency—the first lab could be eliminated before deploying enough of its custom silicon. So frontier players actually need to deploy the same platforms, because they're competing at the model level, not the chip level. That's why independent frontier chip companies exist long-term.

— Walter Goodwin

In their own words · checked verbatim

Right now we're trying to build a single chip. If you crack open an NVIDIA system, it has anywhere between kind of six and nine custom chips, all built by NVIDIA, to come together to build something that is really, really potent. So there's already a gap there.

Walter Goodwin0:00

we need to find a way to pour more compute into these models at test time

Walter Goodwin1:01

one of the things that I think is very striking is there's a lot of relatively identikit chips out there.

Walter Goodwin3:02

when Henry Ford sort of asked the rhetorical question of what people would have said they wanted and, you know, faster horses rather than a car. The snappier chatbot is kind of the faster horses of kind of fast inference.

Walter Goodwin12:25

if you can just find a way to structurally carve out a three to six months advantage, you will be winning all of those deployments.

Walter Goodwin22:12

We've scaled flops like a million fold in the last 20 years. Memory bandwidth has gone up about 40x in the same time frame.

Walter Goodwin30:32

I could die in the nine months before I get to deploy enough of that chip that I've also now gained that kind of 5x in computational efficiency.

Walter Goodwin34:36

Figures

Custom chips in NVIDIA systems6 to 921:09
Broadcom market cap$2 trillion3:02
Fractile headcount~1508:19
Tape-out cycle3 to 5 months18:42
Chip amortization payback period3 to 5 years18:42
Fractile bandwidth advantage vs. HBM25x28:28
Memory bandwidth growth over 20 years~40x30:32

Glossary

HBM
High-bandwidth memory co-packaged with GPU/TPU dies; the standard memory for mainstream AI chips.
ASIC house
Chip foundry backend partner (like Broadcom) that converts large companies' chip designs into tape-out-ready layouts.
GDS2
Layout file format: the final chip design file submitted to wafer fabs, recording the physical position of every transistor and metal layer.
DRC/LVS clean
Design rule and layout-vs-schematic verification: layout passes fab design rules and matches the schematic—the final gate before tape-out.
MFU
Model FLOPs Utilization; the ratio of actual useful compute on a chip to its theoretical peak FLOPS.
RSI
Recursive self-improvement: AI using its own capabilities to iteratively enhance development processes. Here, a metaphor for chip design automation.

How to listen

Who it's for

Engineers and investors tracking AI inference chip roadmaps, compute procurement, and semiconductor startup ecosystem dynamics.