When AI fails, the only thing America can demand is: shut down entire data centers
Verification tech to disprove safety claims doesn't exist yet, and formal national certification would take years even if it did. When crisis hits, the only enforceable demand is the most extreme one: shut down entire data centers. There's no middle ground.
The video won't play here. Listen to the audio instead:
The argument · timestamps estimated from transcript position
In crisis, the only demand America can make is: shut down entire data centers
The Harris brothers' logic: weak US-China trust combined with rapid AI progress means when a serious incident occurs, the only way to verify the other side's compliance is to demand shutdown of a cluster large enough for satellites to detect. But satellites only see thermal signatures—crude data. This demand destroys tens of billions in assets; GPU depreciation dominates the operating budget. The actual work: develop verification tech that replaces shutdown—e.g., only require half down, verify the other half, and save those billions.
— Ed HarrisIntelligence agencies take years to certify verification tech for national use
Even if verification tech exists now, intelligence agencies won't adopt it immediately. Historically, this type of capability takes years from development to formal certification as a ‘National Technical Means’ (NTM). Harris recommends verification startups contact intelligence agencies now to align research direction, so when crisis hits they don't evaluate from scratch. They can decide faster which verification method to deploy, rather than jump straight to the most expensive option: full shutdown.
— Jeremy HarrisTrue commitment shows in willingness to expose your own failures
More convincing than a carefully worded speech: whether China's own regulators report system failures to US frontier labs at equal granularity—e.g., a model that passed ‘ideology alignment’ in testing but drifted after launch. This level of self-exposure would signal real belief that ‘we're both training something we don't fully understand’ is a shared risk, not just diplomatic posturing.
— Ed HarrisEnterprise cloud giants charge 2–3x more for GPUs—the premium isn't the chip
Silicon Data research: Amazon and Microsoft GPUs cost 2–3x more than emerging GPU-native cloud providers, often higher. Steve Ho's explanation: this premium isn't the chip. It's bundled compliance, security analysis, and long-term relationships enterprises depend on—like the same hamburger costing different amounts at a street stall versus high-end steakhouse. The premium can't be parsed as simple arbitrage, so these vendors form a separate non-arbitrageable category.
— Steve HoRising token prices reflect users switching to pricier models, not intelligence getting expensive
Silicon Data's token price index recently rebounded, but Steve Ho clarifies this doesn't mean intelligence is getting more expensive—frontier model token prices have actually fallen continuously since late June, from ~$4 to $1.6, a 50%+ decline. The index rebound reflects user behavior shifting: more willingness to pay for flagship models. The index is ‘spending-weighted’—like car-price averages rising because buyers shift to luxury models instead of base models. Unit costs keep falling; weighted average rises.
— Steve HoThe human expert found the genetic deletion first; AI only replicated the finding later
Daniel McKinnon's son Owen died of a rare lung disease. The original genome sequencing report: undiagnosed. A human expert later found the missed cause: a 91-kilobase deletion removing an enhancer regulating a gene, located about one million bases away from the gene—far beyond the standard filter most labs apply (‘only check within 1kb of gene boundary’). Daniel later wrote an AI analysis pipeline that re-discovered the same variant. Sequence matters: human found it first. AI replicated, not diagnosed ahead of all human experts.
— Daniel McKinnonSpecialized genetics tools hit 10%; general models hit 50% right out of the box
Daniel's company Gamow Labs maintains Rarebench, a benchmark for variant pathogenicity ranking. Best traditional ML tool, Lyrica, scores ~10%. Unoptimized Opus 5.5 run raw in Claude Code: ~50%. Starker still: earliest o3 model scored 0% on this benchmark. This gap informed Daniel's founding thesis: don't solve this by squeezing specialized small models; solve it with frontier general models plus good harness design.
— Daniel McKinnonThe cost of slowing AI is that half these families never get answers
Nathan concedes the real cost of supporting an ‘AI pause’: Gamow's current model scores ~50% on Rarebench, leaving half of rare-disease cases unresolved. He still believes safety compromises are worth it—frontier labs won't lose business momentum—but stresses it's not free. For a specific family, it's a concrete, high-cost trade-off.
— Nathan LabenzIn their own words · checked verbatim
Certain kinds of transparency can be stabilizing. So the the kind of transparency that goes like, hey. We're giving you enough vision into what we're doing to see that we are not doing the thing you fear most.
Ed Harris0:21
data centers put off a hell of an energy footprint. You know, the thermals on those are really, really bright. Data centers are huge.
Jeremy Harris5:57
So how long did it take similar technologies in the past to get used, to get vetted, and to become what's known as national technical means, NTMs? Right? And the answer is is years, depending on the technology, but very often years.
Jeremy Harris10:31
So in the case of hyperscalers, indeed, you are observing correctly. They charge regularly, consistently, at least two to three times, sometimes more compared to a typical a new cloud.
Steve Ho34:31
since, like, maybe late June through, like, basically, you know, middle of this month, it has been on a very sharp downward trend of going from some $4 to, what, you know, 1.6 or something that's more than 50% drop.
Steve Ho39:31
this enhancer is a megabase 1,000,000 bases up upstream from the gene.
Daniel McKinnon1:01:22
the best performing, like, I'd say, machine learning based approach in terms of variant ranking is called Lyrica. I think it scores something like 10% on our benchmark. And right now, Opus 5.5 in Cloud Code, just a vanilla thing, scores something like 50%.
Daniel McKinnon1:08:07
I don't think the frontier model companies are at risk of, like, not being able to grow business, not being able to grow revenue, becoming surpassed if they pace their frontier efforts.
Nathan Labenz1:22:42
Figures
| GPU price premium: enterprise vs emerging clouds | 2–3x, sometimes higher | 34:31 |
| Frontier model token-price decline (late June–now) | From ~$4 to $1.6, over 50% drop | 39:31 |
| ACMG score to classify variant as likely pathogenic | 6 or higher | 1:04:16 |
| ACMG score gain from one functional validation experiment | 2–4 points | 1:06:33 |
| Rarebench score: Lyrica vs Opus 5.5 raw | 10% vs 50% | 1:08:07 |
| o3 model's initial Rarebench score | 0% | 1:18:14 |
| Joel Borgen's editing deletion rate on AI draft | ~14% | 52:11 |
Glossary
- Variant of Uncertain Significance (VUS)
- A genetic variant whose pathogenic effect is unknown—neither clearly benign nor disease-causing, existing in uncertain middle ground.
- ACMG classification
- A standard scoring system used clinically to grade the pathogenicity of genetic variants.
- National Technical Means (NTM)
- A verification capability formally certified by intelligence agencies for compliance checking.
- Enhancer
- A DNA region distant from a gene's body that regulates its expression level.
- Neo-cloud (emerging GPU cloud provider)
- A specialist GPU rental service, distinct from large established cloud vendors.
How to listen
Founders and investors concerned with US-China AI competition mechanics, cloud GPU arbitrage opportunities, or real-world AI diagnostic performance in specialty domains.
If AI's role in fiction writing doesn't interest you, skip 47:30–58:10.