The world is too loud. Read what matters.

80,000 Hours Podcast

The window to slow down has already closed; superintelligence arrives in two to three years

Geoffrey Irving, the UK's former head of AI safety science, argues that full superintelligence is two to three years away and the window for slowing down has already passed; the hope for alignment lies in scalable oversight, character research and theoretical breakthroughs, not in precisely specifying a utility function.

AI alignmentSuperintelligenceSlowing downScalable oversightAISIAGI

The video won't play here. Listen to the audio instead:

This episode takes on the central arguments in AI alignment head-on: whether slowing down is feasible, when superintelligence arrives, and where lab strategies fall short. It delivers a large number of actionable judgments at high information density.

The argument · timestamps estimated from transcript position

0:00

Fewer than 10 people nodding could slow AI down

Irving states his position plainly at the outset: the time to slow down is now, or even earlier. If you work through carefully when we ought to slow down, he explains, the answer comes out as ‘some point in the past’, because we are too close to that crazy future. He argues that agreement among fewer than 10 key decision-makers — lab CEOs, the relevant people in China, a few heads of state — would be enough to make slowing down happen; what is hard is doing it unilaterally. This judgment sets the tone for the whole conversation, and it explains why he sees alignment as requiring global coordination rather than a purely technical fix.

— Geoffrey Irving
3:07

A one-week training run does not rule out multi-year plans

Responding to Rohin Shah's view that training runs are short, so models will not learn to take over the world, Irving argues this conflates the time scale of an overall plan with the time scale of the components of a task. Even if every subtask completes within a week, combining multiple subtasks can add up to a multi-year plan. That means training duration is not a safety guarantee: a model can execute a long-horizon strategy by decomposing the task, without having to complete the full chain of reasoning inside a single training process.

— Geoffrey Irving
9:38

The three big labs are running the same alignment playbook

Irving sees the strategies at OpenAI, Anthropic and DeepMind as the same combination: character training plus scalable oversight plus monitoring. He contrasts their approaches to character training: Anthropic leans toward virtue ethics, OpenAI has more rules and leans more corrective, and DeepMind's approach is less clear. Whether any of this works at the level of superintelligence has no strong argument behind it today. The observation reveals that alignment practice at frontier labs lacks theoretical grounding and is largely empirical tuning.

— Geoffrey Irving
21:09

Show the evidence, hide the counterevidence, and the judge is fooled

Irving names a key obstacle for scalable oversight — ‘ambiguous arguments’: a model may present only the evidence supporting a claim while concealing the counterevidence, precisely because the counterevidence is the decisive part. He cites Beth Barnes's experiments at OpenAI, where human debaters fooled judges with complicated but wrong arguments, because neither side could find the flaw. This obstacle has no good theoretical solution, which suggests that oversight resting purely on human feedback may fail against smarter models.

— Geoffrey Irving
29:34

The claim that models only excel at verifiable tasks is overrated

Irving's modal expectation is full superintelligence within two to three years. The main uncertainty is whether models can get good at fuzzy tasks — intuition, vague planning. He thinks people overrate the possibility that models will only ever be good at verifiable tasks, and believes that through the data flywheel, models will gradually extend into more non-verifiable tasks. This timeline judgment implies the window for slowing down and for safety research is extremely short, and it supports his opening claim that the time to slow down is now.

— Geoffrey Irving
45:09

Government AI safety work gets its real leverage from policy

Irving lays out AISI's future priorities. First, misuse risk — dangerous capability evaluations, which require combining strong research capability with national security expertise. Second, mitigations, covering both misuse and loss of control; AISI already has a strong safeguards team. Third, policy, which was Irving's main reason for joining AISI. As the largest government AI safety research institution, AISI's policy advice should be grounded in tactical reality. These three items sketch the role of government in AI safety: evaluation, defense, and policymaking.

— Geoffrey Irving
1:21:36

Alignment is not setting a goal, it is designing the training process

Irving argues that before superintelligence arrives, we will not be able to precisely specify its utility function. The only hope is that the training process has error-correcting properties, so that we land in some ‘basin’ rather than hitting a precise target. He offers one concrete idea: a branching diagram of training space, using 1000 checkpoints to display a model's training trajectory, so that different algorithms can be tested and understanding of branching behavior can grow. This is an attempt to shift alignment from setting goals to designing training processes.

— Geoffrey Irving
1:38:14

RL mainly works by selecting against ineffective reasoning patterns

Irving sees RL as doing two main things: selecting against ineffective reasoning patterns, and compensating for the fact that humans usually write down only the final answer, which forces models to reason in latent space. The arrival of o1 was a combination of tuning RL algorithms and improvements in base model capability. He predicts the frontier of alignment technique will be the combination of scalable oversight plus character plus learning dynamics, but that organizations will pursue diversified strategies rather than bet everything on one approach.

— Geoffrey Irving

In their own words · checked verbatim

Now. Now. If we were to carefully analyse this question of exactly when we should slow down, it would be like a while ago in the past, because we’re just too close to this crazy future.

Geoffrey Irving0:00

On factual questions it would be not racist; it was very happy to write incredibly horrible poetry about racism.

Geoffrey Irving14:31

One of the winning strategies was basically a debater would produce a very complicated argument that sort of sounded true, was false, but neither of the debaters — not the liar or the honest debater — knew where the flaw was.

Geoffrey Irving23:31

some degree of your conscious train of thought is a hallucination as you go along the world.

Geoffrey Irving56:12

"Rationalisation" there is a pejorative word, but it’s also just intrinsically how intelligence works, even for humans.

Geoffrey Irving1:06:35

Hope is not like a lot of probability. Just like, we should try.

Geoffrey Irving1:23:51

If you had the other LBJ that was opposed to the first one, that's saying, "By the way everyone, you realise what he's doing here? He's cheating the vote system." And everyone is like, "That's ridiculous. That's clearly unfair. Let's fix the rule to break that." I think that intervention would get you so much power over the misaligned components of LBJ that I think it's within hope to imagine getting that story right.

Geoffrey Irving1:28:54

I told Dario — this is annoying — that I would write a document called "Language is enough to get to AGI" — and then I didn't write it until 2019. So early 2019 is when I actually wrote the document.

Geoffrey Irving1:36:46

Figures

Key people needed to slow downfewer than 100:00
Rules in Anthropic's constitution510:46
Acceptable performance cost of a temporary pause2-10x slower34:14
Time until superintelligence (modal expectation)2-3 years29:34
History of the alignment fieldat most 20-25 years1:18:26
Age of empirical character research1-2 years1:19:30
Time until nanotechnology and uploadsa few years to twenty years59:28
Compute to solve Pentago10^17 or 10^18 FLOPS1:31:00
Year he began arguing for the RL+LLM path20181:34:05
o1 release20241:36:46

Glossary

scalable oversight
Oversight methods that keep behavior in line with intent even once AI capability exceeds that of the human supervisors.
ambiguous arguments
A mode of reasoning that offers only supporting evidence while hiding decisive counterevidence, and can therefore deceive a supervisor.
character training
Shaping a model's personality traits through training in order to reduce harm; part of current lab alignment strategies.
effective field theory
A theory that ignores microscopic detail and describes physical phenomena with macroscopic parameters; Irving uses it as an analogy for alignment theory.
data flywheel
The loop in which a model generates data used to train the next generation, continually extending the range of capability.

How to listen

Who it's for

Worth your time: alignment researchers, leads of frontier model training teams, AI policy practitioners, and any technical decision-maker whose plans are sensitive to superintelligence timelines.

Skip

The compute estimate for solving Pentago around 1:31:00 can be skipped; it does not affect the main argument.