The US-China AI gap is six to eight months: complete distillation bans won't slow much
Epoch AI researcher JS Denain argues the leading-model capability gap between the US and China is only six to eight months, and even under a complete distillation ban, the gap would expand merely from six months to eight or nine months within half a year.
The video won't play here. Listen to the audio instead:
The argument · tap a timestamp to hear it
The public evidence for RSI is actually thin
OpenAI and Anthropic both published blog posts on how AI accelerates AI research, but JS Denain argues this evidence doesn't sufficiently support claims of ‘imminent self-reinforcing capability acceleration’ or ‘full automation of AI researcher roles’. The most striking signal is research staff spending on Codex growing roughly 2x monthly, but Denain sees this as adoption driven by improved model availability rather than proof of imminent software-level intelligence explosion. He stresses such metrics are inherently fuzzy—it could simply be traffic that previously scattered across ChatGPT now consolidated under Codex.
— JS DenainAI can't yet drive your research agenda
OpenAI's own acceleration blog shows researchers using Codex more for engineering and debugging tasks, but neither interviews nor conversation logs reveal material gains in high-level strategic thinking—resource allocation, research-direction selection. Nathan Lambert clarifies the distinction with an analogy: a PhD student's raw output might genuinely accelerate tenfold, but the advisor's entire research agenda won't necessarily accelerate tenfold; organizational bottlenecks absorb individual productivity gains.
— Nathan LambertRobot industry boom: underestimated supply elasticity
Nathan Lambert argues the robotics path will resemble autonomous driving, not the LLM boom—factories are hard to build, long-tail reliability problems abound, meaningful robotics scale is invisible for two to five years. JS Denain counters: if robots can replace blue-collar workers, the addressable market is enormous; US data-center build speed shows financial incentives overcome supply constraints, and US underinvestment in some sectors looks demand-constrained, not capability-constrained. He concedes his own knowledge of robot-capability progress is limited.
— JS DenainUS-China gap is six to eight months; distillation won't move it much
Across multiple estimation methods, JS Denain places the gap between US and Chinese leading public models at six to eight months (release date to release date). Even under the extreme scenario of ‘complete distillation ban’, he estimates the gap would expand only from six months to eight or nine months six months later—not the twelve months some expect. He emphasizes the raw-compute gap is orders of magnitude larger, yet the capability gap is much narrower; this gap itself demands explanation. Distillation is one candidate explanation, not the sole one.
— JS DenainWhy he came to believe distillation actually works
Denain acknowledges underestimating distillation's impact previously. Evidence that shifted his view includes the Stolen Thoughts paper—showing model reasoning traces are relatively easy to extract via jailbroken API access—and Claude's router exposing actual prompt distributions, far richer than synthetic benchmark samples. He also mentions claims on another podcast that ‘post-training in China already achieves 80% of results’, which led him to reweight the importance of SFT-based language-trajectory training, not RL exploration alone.
— JS DenainProving something possible is the biggest assist
Beyond distillation, Nathan Lambert proposes an alternative explanation for the shrinking China-US gap: the ‘four-minute mile effect’—once OpenAI and Anthropic prove a technical path works, Chinese labs don't waste resources exploring dead-ends; they replicate the shortcut instead. JS Denain accepts this logic but cites counterexamples like reasoning models (DeepSeek Math predated o1), which don't fit the pattern neatly. Both agree this looks more like organizational and research-timeline differences.
— JS DenainModel safeguards can't stop it; even OpenAI was breached
Just 24 hours before recording, someone used a public Claude model to breach OpenAI and gain internal access. Denain says it's hard to update whether this signals ‘insufficient model guardrails’ or ‘insufficient corporate security’, but the reality is stark: open-source models have weaker guardrails but weaker capabilities; closed-source models have slightly stronger guardrails but vastly more power. The net safety impact is opaque. Both the Stolen Thoughts paper and the OpenAI incident show that even under strong company incentives to prevent it, safeguards get circumvented.
— JS DenainPost-training is the industry's biggest black box
Denain sketches his mental model of leading labs' post-training workflows: a large post-training team split into domain sub-teams (math, biology, finance), each managing data, RL environments, and hyperparameter tuning, with results later packaged and competing in one large RL training run. Nathan Lambert agrees with the organizational breakdown but points out the real complexity gap: whether it's ‘multi-expert distillation merging’ or ‘sequential large-scale RL’—a choice that might itself bottleneck progress speed. Frustratingly, almost no external visibility into actual processes exists.
— JS DenainIn their own words · checked verbatim
that plot, for example, just shows a 2X a month increase in Codex spending by researchers, and that does seem to me to be some evidence of they’re getting a lot of value out of this probably. I don’t think this is strong evidence that in six months we have a software intelligence explosion.
JS Denain0:00
I think an analogy I have is, I think that junior PhD students will be 10X as productive, but from the advisor’s perspective, their research agenda will not proceed 10X as fast.
Nathan Lambert9:38
if you look at data centers in the US, there has been extremely fast build-out. If you had robots that were just literally able to substitute for blue-collar human workers, I feel like the financial incentives would be huge.
JS Denain20:35
so currently if I’m saying six months gap, then in six months, it’s probably not literally 12 months, right? So, my guess is, yeah, you are at like eight or nine, something like that. You go from like six to nine maybe.
JS Denain27:39
I think the Claude routers just give you a bunch of useful data to train on, and that will generally improve capabilities. I think that’s one explanation.
JS Denain29:50
four-minute mile style effects of like, you show that something is possible at all and then people have the conviction to go down that route, don’t waste a ton of resources exploring a bunch of other things, is the thing you’re pointing at here.
JS Denain38:27
there’s definitely clearly examples like this one, or really the Stolen Thoughts paper or something. It just seems like the safeguards weren’t good enough to prevent that from happening, even though there’s strong incentives for the companies to prevent it.
JS Denain48:46
Figures
| Codex spending growth rate | ~2x monthly among researchers | 0:00 |
| Evan Hubinger's existential risk probability | at least 10% | 4:09 |
| US-China top open-model capability gap | 6–8 months (release date to release date) | 24:47 |
| Projected gap under complete distillation ban | expands from ~6 months to 8–9 months | 27:39 |
Glossary
- RSI
- Recursive self-improvement: an AI system accelerating AI research itself in a self-reinforcing cycle
- ECI
- Epoch Capability Index: Epoch AI's composite metric for measuring and comparing model capability levels
- distillation
- Training a weak model on stronger model outputs to quickly acquire partial capabilities
- four-minute mile effect
- Once someone proves something possible, later entrants need not explore dead-ends—they replicate the shortcut instead
- jaggedness
- AI capability progressing unevenly across different tasks, strength and weakness alternating
- Goodharting
- Once a metric becomes an optimization target, it ceases to reflect reality
How to listen
Entrepreneurs and investors tracking the US-China AI gap and RSI risks, particularly those seeking concrete mechanisms rather than broad conclusions.
41:22–47:35 on recruitment-data research methodology—skippable.