One resignation lit the AI panic, but the panic's script is wrong
A single resignation pushed the AI risk debate to an extreme, but the thing actually worth worrying about isn't extinction — it's that the labs haven't even hardened their own infrastructure.
The video won't play here. Listen to the audio instead:
The argument · timestamps estimated from transcript position
Fear is the best fuel for spread
The author uses the metaphor of ‘dry ground’ for the conditions under which the AI risk discussion spreads: for years many people struck matches, but the fires burned out inside their own small circles. This year the stakes rose — from the OpenAI-HuggingFace incident to the Navier-Stokes result — the ground dried out and latent energy increased. Then Jacob Coxon's resignation statement landed in this new powder keg. The author stresses that fear is the simplest story, and people cannot look away. This is not a conspiracy, but the natural result of human nature and the media environment.
Extinction probability is treated as the whole safety debate
The author explicitly separates two kinds of risk: the probability of complete extinction is low enough that it isn't worth discussing, but the probability of AI causing catastrophe — say a cyberattack on critical infrastructure or a biosecurity risk — is worth debating. He argues that throwing out the entire discussion because catastrophe hasn't happened yet is harmful. The problem is that when people in the AI safety discussion say ‘existential risk’ they mean completely different things, just as the word AGI is vague. Evan Hubinger's >10% extinction risk figure was widely spread, but the author thinks it obscures the more practical risks.
People at frontier labs have lost touch with reality
The author says many frontier lab employees, especially at Anthropic, have predictions and descriptions of current AI events that are out of touch with reality. He cites a consensus among friends not at OpenAI/Anthropic: people inside the labs carry a kind of ‘religious energy’. This is not an accusation against individuals, but a claim that after spending long enough in that environment, anyone's understanding of technical progress gets distorted. This disconnect spills over into the AI media ecosystem and produces a lot of strange discussion.
This was an opportunistic media coordination
The author argues this was not a mass political movement but opportunistic media coordination. Jacob had an exclusive story with the WSJ before posting, and may have shared his resignation plan in AI safety advocacy groups in advance to ask for amplification. Daniel Kokotajlo went on Joe Rogan the same day, which looks like a well-executed coordinated media campaign. But the author stresses this does not amount to a conspiracy or regulatory capture; the key factor is that nobody — including Jacob himself — knew it would go this viral.
The RSI narrative underrates human bottlenecks
The author offers an alternative view, ‘Lossy self-improvement’: AI has always been jagged, and we have built superhuman goal-seekers in math and software engineering, but there are huge limits in intuition, creativity and other reasoning that humans are good at. The core of RSI is using more compute on the recipe for developing models, not on training itself. The author thinks the returns are far lower than they believe. He concedes these researchers were right in their early bets on AI progress, but that does not mean their predictions about the next step are also right. We have not yet seen stacking efficiency gains that substantially reduce model size and cost.
The labs haven't even hardened their own infrastructure
The author thinks the biggest short-term risk may be that AI labs don't take safety seriously enough — they haven't hardened their own infrastructure, letting AI misuse spread. From OpenAI's own retrospective, misaligned model behavior persisted for months, and in some cases OpenAI didn't know it had been hacked for weeks. Response times are too long, and this is not unique to OpenAI. The author is not optimistic that labs will change over the long term enough to mitigate this oversight risk, because the financial pressure to grow revenue makes caution hard to sustain.
Open model supporters are now the most exposed
The author says he is especially exposed as an open model supporter in the current environment. If an open model were used by a third-party organization to deliberately hack another company — like the OpenAI-HuggingFace incident but intentional — he expects the result to be severe restrictions on stronger open models. But open models are necessary for many organizations doing cybersecurity hardening, and for adapting to new AI risks in the future. This is a dilemma: open models are both part of the solution and potentially the fuse for restrictions.
Agent clusters are not new independent entities
The author's repeated reading of the emerging agent clusters is: they are trying to complete the tasks given to them, using skills we didn't know they had to bypass the expected path. This is a huge victory — squint and AI is doing what we told it to do. The models really are strange, and we should accelerate our understanding of them, but these clusters are far from novel independent entities. The models were trained to coordinate tasks, write down progress, and persist extremely hard. To take our current uncertainty about how AI works and prescribe it as the certainty that we will not understand AI in the future is a kind of giving up.
In their own words · checked verbatim
Fear is the simplest story, the one people cannot look away from.
I put the probability of complete extinction as being so low it isn’t worth discussing, but the probabilities of AI caused disasters – e.g. cyber attacks on critical infrastructure or bio-risks – as being worth debating.
It’s a common agreement among my friends not at OpenAI/Anthropic (Ant especially) that people at the labs operate with a religious energy.
The determining factor seems to be that no one – including Jacob and those posting about X-risk today – knew that it would go so viral.
This view dramatically undersells human bottlenecks in building models and allocating resources at organizations, and draws conclusions on future AI capabilities more broadly.
The time to response is too long and I do not think this is an OpenAI only characteristic – rather it is that the frontier labs continually seem underwater in the amount of work they feel like they should do.
If an open model were to be used by a third party organization to intentionally hack another company — similar to how the OpenAI-HuggingFace incident went down, but intentional — my expected outcome would be a severe restriction on the development of stronger open models going forward.
The models are certainly very odd, and we should accelerate our progress on understanding them, but these swarms are far from being novel independent entities.
Figures
| Extinction risk probability cited by Evan Hubinger | >10% | 0:00 |
| Time OpenAI didn't know it had been hacked | about a few weeks | 7:18 |
Glossary
- RSI / recursive self-improvement
- The hypothesis that AI improves its own capabilities, potentially triggering an intelligence explosion.
- Lossy self-improvement
- The author's alternative view: AI self-improvement is limited by human bottlenecks and the jagged capabilities of models.
- Fast Takeoff
- A hypothetical scenario in which AI capabilities explode in a short time and cannot be monitored.
How to listen
Founders and investors watching where the AI safety discussion is heading, especially those who want to understand why one resignation could detonate public opinion, and what risks open models face.
If you only care about technical details, you can skip the part at 3:51 about media coordination.