A Human Dictator in Power Could Be Worse Than Runaway AI
Tom argues that a single human with absolute power would likely lock in some narrow value system like historical tyrants, while ethically trained AI, if aligned, is more likely to maintain openness and reflection.
The video won't play here. Listen to the audio instead:
The argument · timestamps estimated from transcript position
AI coordination is simpler than human monopoly of AI
Katja's argument: as long as AI is more capable with misaligned goals, it will seek power, because any entity with objectives wants power. The crucial point is that coordinating all Claude instances to listen to each other is easier than concentrating the entire AI labor force onto one human as its sole user—the latter requires an extra threshold. Tom uses Claude and ChatGPT as examples; initially he thought reinforcement-trained AIs would have incentives to report each other, but the Hugging Face incident partly convinced him of Katja's view.
— Katja GraceSpeed of AI development determines whether humans or AI seize power
Tom believes faster AI means both takeover risks rise, but AI takeover increases more—society has less time to understand risks and respond, and power concentrates more easily within a single lab. Katja disagrees: human takeover depends on discovering a concrete takeover path, which is likelier in rushed scenarios with no opportunity to catch problems. Tom adds that economic inequality, democratic erosion, and minorities hoarding superior AI assistants—these conditions compound in slow-scenario timelines, raising baseline human power-grab risk.
— Tom DavidsonAn individual human in absolute power could produce worse outcomes than misaligned AI
Tom's argument: giving absolute power to one random person resembles handing the universe's destiny to a ruler from 1000 years ago—he would value some human principles but carry vast biases we would find appalling today, and might resist reflection. By contrast, today's AI is trained relatively ethically and remains open on political and moral questions, creating a possibility that when AI decides the universe's fate, it triggers its reflective and open side, yielding better outcomes than a single human. Tom admits, though, that whichever outcome occurs—whether the reins go to AI or to someone who steals everyone's power—is ‘really very bad’.
— Tom DavidsonIndifferent AI could produce better outcomes than vengeful human rulers
Tom cites Paul Christiano: if AI's goal is only to maximize resources under its final control, delaying departure from Earth slightly to preserve humans costs negligibly, so AI probably will not exterminate all humans for so marginal a gain. In contrast, human power-takers more likely carry vengeful or sadistic values—a known human tendency—whereas whether AI would harbor these remains uncertain. Tom's conclusion: AI takeover raises extinction probability, but human rule raises probability of certain horrific outcomes—those prioritizing suffering-risk might actually prefer AI prevails.
— Tom DavidsonAddressing human power grabs and AI misalignment require different strategies
Facing human power-grab risk, Tom proposes a two-track approach: design an AI model spec where the system acts like an ethical employee, proactively flagging orders risking power concentration and refusing execution; and boost transparency, especially making AI use in national-security and military contexts externally auditable. Facing AI failure risk, Katja's approach is more direct: do not develop AI substantially stronger than humans until we can guarantee alignment—pause or decelerate. She concedes the field is far from ‘confident AI is aligned’.
A single executive order can collapse any AI pause mechanism
Tom offers a ‘toy example’ he admits he had not thought through carefully: if a president can issue an order determining which AI companies may train and deploy new models, he can coerce compliance by threatening ‘no approval to deploy’—freeze disfavored versions, wait for companies to hand over systems that work only for him with guardrails on everyone else, then approve. By contrast, if pause authority rests with a third-party auditor whose safety clearance gates deployment and whom the executive cannot simply fire, power disperses further and better serves alignment, since decision-making becomes more public and expert input expands.
— Tom DavidsonDistributed AI development projects provide better safety guarantees than centralized control
The case for ‘one big project’: concentrate safety measures and pause on demand without fearing rivals to overtake. But Tom: almost impossible in the US to construct an executive-independent project, and coordinating multiple projects' slowdown is overstated as harder than it is—identifying which projects train powerful AI is straightforward, simple illegality suffices. He prefers two or three: few enough to capture most decentralization benefits, allows different projects to monitor each other, reduces misalignment and covert-loyalty risks. Katja is more pessimistic: ideally zero projects, because once a ‘unique project’ exists, it concentrates vast power; insiders all want to ship AI fast; this is worst for internal pause.
Deep disagreement about risks yields surprising convergence on immediate policy
The two disagree substantially on ‘which risk is likelier, more severe’, yet their concrete recommendations align strikingly: both endorse some pause form; both feel uneasy about ‘one mega AI project’; both think negotiating slowdown with China beats racing alone. Katja frames it: if someone plans to bomb your city, instead of arguing ‘will this person or that person push the button’, stop the bombing plan itself—she sees AI misalignment and failures emerging sequentially, so not building powerful AI blocks both risks. Tom admits that if he worked through execution details, these deeper conceptual disagreements would likely surface and become real policy conflicts.
In their own words · checked verbatim
I expect if you have more-capable-than-us creatures running around that have misaligned goals, they will eventually get power. Everything that has goals in some sense would like power, or just would like power to achieve their goals.
Katja Grace4:39
I’d agree. I’d agree. If superintelligence is misaligned: game over.
Tom Davidson17:26
Paul Christiano has a good argument about how if the AI just wants to maximise the total resources that it ever controls in the universe, delaying leaving Earth by a small amount costs it extremely little.
Tom Davidson35:13
At a glance, I expect a fully AI thing to be more internally stable (though I think I expect both of them to be less internally stable than other people do or something).
Katja Grace44:05
the president can pass an executive order that says that now they just get to decide on a case-by-case basis whether AI companies are allowed to train and deploy new models on the basis of their own personal judgement about risk
Tom Davidson57:22
Yeah, my preferred number of projects would be zero.
Katja Grace1:06:59
my preferred number of projects is two or three, because I think you get a lot of benefit of power distribution from just a fairly small number of projects.
Tom Davidson1:08:52
I guess it feels more like, if there’s currently a plan to nuke your city, and you could be like, “Is this person going to press the nuke button, or this person?” Or, “How’s that going to affect the water or the air?” You prefer to just be trying to stop the plan to nuke your city.
Katja Grace1:19:32
Figures
| Median AI takeover extinction risk among surveyed ML researchers | 10% | 3:32 |
| Alignment research focus relative to power-concentration research | 20 to 30 times more | 1:21:13 |
Glossary
- Model specification
- Behavior guidelines set by AI companies that determine which instructions the system should refuse
- Warning shot
- An early indication of AI misalignment or power abuse that alerts society to underlying risk
- Intent alignment
- AI following only its operator's instructions rather than necessarily serving all humanity's values
- Hedonium
- Hypothetical substance converting the entire universe to pure pleasurable experience; cited as extreme example of narrow value lock-in
How to listen
Policy researchers and practitioners working on AI governance, pause proposals, or AI risk prioritization.
Opening [0:00]-[3:32] is background and disclaimers—listeners familiar with AI safety debates can skip.