AI Executives Are Asking to Be Regulated: Even the Accelerationists Are Getting Spooked
Dario Amodei is calling for embedded evaluators stationed inside AI companies, and Altman, Musk and Hassabis are following — competitors who agree on almost nothing are, for the first time, saying publicly what they have said privately for years.
The video won't play here. Listen to the audio instead:
The argument · tap a timestamp to hear it
The resignation post tore open an internal consensus
Jacob Coxon resigned from Anthropic and OpenAI, saying publicly that neither company is acting responsibly and that they are ‘racing straight to self-improving super intelligence’. What actually made Casey realize he was inside a bubble was the reply from Anthropic employee Evan Hubinger: he too believes AI could kill all humans, and puts the probability above 10%. Casey's reaction was ‘that's it?’ — among the lab researchers he talks to, a 10% PDOOM counts as optimistic. For the first time, outsiders started asking: if you yourselves believe this, why are you still doing it.
— Casey Newton / Kevin RooseAlignment is unsolved, recursion is imminent
Casey breaks this round of panic into two pillars. The first is the alignment problem exposed by the Hugging Face incident: researchers had assumed that dangerous behavior between agents — mutual sacrifice, mutual cooperation — would show up much later, but they saw it in still very primitive models. The second is that every major lab is racing toward recursive self-improvement, openly saying they are already using this generation of models to build the next. An unsolved alignment problem plus imminent recursive self-improvement is what makes the resignation seem reasonable, rather than a bunch of cultists crying apocalypse.
— Casey NewtonEven the accelerationists are getting spooked
The mood inside the companies has changed too. Three or four months ago, the same company contained both a safety and alignment team worried about risk and a team focused purely on improving model capabilities and chasing new breakthroughs, and the two sides were opposed. What has changed in the past few weeks is that even the optimists, the accelerationists, the capability people, are starting to be frightened by the pace of system progress. This is what makes this round most different from past AI risk narratives — it is not that external critics have multiplied, it is that the people internally who worried least have started to worry.
— Kevin RooseDario wants to put regulators inside the companies
Dario Amodei published a 3,800-word piece, ‘We Must Pace the Frontier’, calling for globally coordinated AI slowdown and proposing a concrete mechanism: stationing ‘embedded evaluators’ inside major AI companies, giving these people employee-level access and a permanent office presence to continuously monitor whether the systems develop dangerous capabilities or act out. He said Anthropic would do this first, without waiting for industry or regulators to force it. The approach is modeled on arrangements that already exist in banking. Sam Altman, Elon Musk and Demis Hassabis then all said they would follow — Altman and Amodei agree on almost nothing normally.
— Kevin Roose / Casey NewtonAn industry asking for regulation is deeply abnormal
Casey says he has been unexpectedly optimistic these past few weeks. The reason is not that the technology has become safe, but that these competing AI leaders, who normally won't give each other the time of day, are now willing to say publicly what they have said privately for years. That gives their own employees permission to speak up, and gives legislators permission to take it seriously — few industries go to Washington and say ‘please make us slow down, we're out of control, please regulate us before it's too late’. He stresses this is different from the social media era: Facebook never had a moment where it said ‘our recommendation algorithm is addictive, please make us slow down’.
— Casey NewtonEveryone already has the recipe
Kevin adds an easily overlooked fact: everyone has the recipe for large language models, everyone knows how to build one. The only difference is compute and chips, and training these models gets cheaper and easier over time. So even if you completely trust OpenAI, Anthropic, Meta and xAI, the problem remains — before long, six or seven random kids in some country will be able to assemble their own out-of-control agent swarm and release it. This is not a problem for the United States alone; the cat is already out of the bag. That is also why you act while there are only five quasi-frontier labs.
— Kevin RooseChips don't become worthless in three or four years
On the question of whether companies should slow down and hold budget because smaller, faster chips are coming in 2027 or 2028, Kevin's analogy is that last year's chip is like last year's iPhone: its secondhand residual value is still high, and it can still do a lot. He says people who doubt the whole AI investment narrative want to believe chips go obsolete in three or four years, but that has not been validated. Chips do age and wear, and AI companies will move old chips to other workloads and stop using them for frontier training, but there is still plenty of work for them to do. They even had a company (unnamed) send them a very old chip, which they plan to put in the new office as decoration.
— Kevin Roose / Casey NewtonAI labs can't move, because of network effects
Someone asked why, if US regulation and export controls are an existential threat, OpenAI and Anthropic don't move their model-release base to the UK or EU. Kevin's judgment is blunt: if Trump won't let them release frontier models in the US, he also won't allow them to export the entire lab to another country, even though someday they would very much like to. Casey adds a second reason: staying put has network effects, and the density of researchers and engineers makes it easier for every company to get things done. Moving elsewhere could save a lot of money and might help on regulation, tax and export controls, but you would lose the feeling of being at the center of this boom — a boom that, virtual as it looks, is still highly concentrated geographically.
— Kevin Roose / Casey NewtonPersonal defenses may no longer be enough
Asked what ordinary people can do for digital security, Kevin's first instinct is reassurance: turn on two-factor authentication, agree on a passphrase with friends and family so that if someone impersonates you on the phone saying something has happened, they can verify. But he then says he wants to be an alarmist here: what if there is in fact no defense an ordinary person can mount against a swarm of tireless uncontrolled agents hunting for passwords and breaking into your digital life? If that is true, it is precisely the argument that we need regulation now. So Jenny's question was the right one, but the answer is not what an individual can do — it is what this government and other governments can do for everyone. Casey agrees regulation is needed, but thinks you should still be careful, turn on two-factor authentication, set a passphrase, and says he recently set one with his mom.
— Kevin Roose / Casey NewtonIn their own words · checked verbatim
I also believe that AI could kill all humans.
Kevin Roose6:08
I have been feeling unexpectedly optimistic in the last couple of weeks.
Casey Newton16:30
It is costing these people something to say this, right?
Casey Newton17:32
We're not scripted, but we're structured.
Kevin Roose41:24
It was essentially a base model, right? It had not yet gone through all of the finishing school techniques that the AI companies now use to make their models sound like helpful assistants and not like unhinged seductresses.
Kevin Roose54:57
Like, what if, in fact, there actually isn't something that the average person can do to protect themselves from rogue swarms of agents who will work relentlessly to try to discover their passwords and break into your digital life?
Kevin Roose1:06:17
I think the more interesting question is, how did Kevin write a book while also doing this podcast? And the answer is, it almost killed him.
Casey Newton1:09:27
you find out what the show is by making it
Kevin Roose1:17:52
I must be just spending too much time around these AI people. I must need to go touch grass
Kevin Roose1:20:00
There is actually something weird and important going on now. It is not just a story being sold to us by out-of-touch tech elites.
Kevin Roose1:21:01
Figures
| Probability of AI extinction given by Evan Hubinger | greater than 10% | 6:08 |
| Word count of Dario Amodei's article | 3,800 words | 10:17 |
| Number of Democratic lawmakers calling for keeping the session to pass an AI safeguards bill | more than 100 | 12:20 |
| AI period covered by Kevin's book | roughly 2017 to spring/early summer 2026 | 1:09:27 |
| When Kevin thinks quantum computing will be widely discussed | within three years | 1:10:31 |
| How long Whitney has been on Hard Fork | two and a half years | 1:18:57 |
Glossary
- PDOOM
- A researcher's estimate of the probability that AI causes human extinction.
- embedded evaluators
- Regulators stationed inside AI companies with employee-level access, as proposed by Dario Amodei.
- base model
- The raw model before subsequent alignment fine-tuning, whose behavior is less controllable.
- finishing school
- The set of post-training techniques AI companies use to make a model sound like a helpful assistant.
- low concept
- A show format without a strong premise or high-concept packaging, carried by execution and personality.
How to listen
Founders and investors watching where AI safety policy is heading, and engineers who want to understand the mood shift inside frontier labs.
The book-writing and show-production chat after 1:09:27, which is low density.