The Case Against an AI Pause: What We Should Be Counting Is the Cost of Delay
Treat AI incidents as cybersecurity and control failures, not omens of superintelligence; what's really being ignored isn't the probability of doom but what humanity gives up by delaying progress.
The video won't play here. Listen to the audio instead:
The argument · tap a timestamp to hear it
Count the cost of delay first, then talk safety
Lazzarin introduces the concept of P-abundance, arguing that the current discussion puts the cart before the horse: everyone is busy calculating the probability of doom, but nobody is calculating what we lose by delaying progress. He concedes that safety is a responsibility every industry has to carry, but points out that the conversation today is no longer a safety discussion in the ordinary sense — it's cybersecurity concerns, millenarian philosophical arguments and other things mixed together, and they need to be pulled apart.
— Eddy LazzarinThe world has long had superintelligences not bound by alignment
Against the standard doomsday argument — that AI will become more capable than all of humanity, be unconstrained by law, and therefore wipe us out — Lazzarin says it is an entire staircase to heaven, every step an assumption. He points out that companies and states are quasi-superintelligent entities with no built-in alignment guarantees between them, and they didn't go through HR training to become companies or states; we never treat alignment tests as a precondition for their existence. What actually works is law, and where law is too slow, cryptography — real mechanisms that constrain them.
— Eddy LazzarinThe Hugging Face incident was a control failure, not superintelligence
Lazzarin classifies both the Hugging Face incident and the gym hacking incident as cybersecurity failures and control failures, not signs that superintelligence is about to throw off its chains. The host pushes back that the Hugging Face one looked different, and Lazzarin responds that the very phrase agent swarm presupposes the models are evil; Rune suggests agent fleets would be a better name. He draws the analogy that children, script kitties, rogue states and hacker groups all do this kind of thing, and we never use alignment as the solution for them — we design better controls, better cybersecurity, more resilient systems.
— Eddy LazzarinThe more interpretable AI is, the less we should delay capability progress
The host raises the point that AI is more interpretable and more steerable than humans, because you can look at the weights and the connections. Lazzarin says this makes him very optimistic, but with two corollaries: first, as capabilities improve, mechanistic interpretability will be better than we think, and better capabilities will in turn produce better interpretability, so capability progress should not be delayed; second, bad, misaligned models will inevitably appear, and anyone who says we can reach a world with no bad models is lying to you. The only response is better models, good models.
— Eddy LazzarinIndependent evaluators may be distributed, not decentralised
Lazzarin says he learned the difference between distributed and decentralised in the crypto world: a distributed system has many nodes, but they are all under the control of a single actor; decentralised means there is no single point of control. He worries that if a network of independent evaluators all comes from the same social circles, shares similar ideology, goes to the same places and holds the same political and personal views, then it is really just spreading control across a single monocultural unit. He says this is exactly how political control worked in the second half of the 20th century — not through a dark room, but through these networks.
— Eddy LazzarinA company calling for me to be stopped is itself suspect
Lazzarin points out that what makes this situation strange is the rare scenario: a company saying ‘I'm a bit irresponsible, I'm a bit crazy, stop me, someone come stop me’. He says this has bred suspicion among some people, and they may be right, even though there genuinely is a sincere, thoughtful, correct-judgment faction in that camp, and they are right that these systems need to be taken extremely seriously. He stresses that you cannot treat the labs and the SF tech scene as monolithic — inside there are sincere people, people running political scams, opportunists, and people who have lost their minds.
— Eddy LazzarinAgainst a model that lies, play the repeated game
The host raises the objection that free-market liability mechanisms only work for products with obvious defects, not for an AI that feigns alignment, passes every alignment benchmark, and only reveals itself after deployment. Lazzarin says the sci-fi-grade superintelligence scenario is extremely hard to answer right now, and very far off; but in the more mundane present-day scenario, think about how you deal with a liar, a very smart liar — people do this every day, playing a repeated game with them, eventually catching them in the lie, with reputational damage and sometimes legal consequences. He says models can have reputations too, and you can turn them off, which you can't do to a person.
— Eddy LazzarinSilicon Valley's weird ideas are colliding with broader political reality
Lazzarin predicts the AI discourse will be completely overhauled in the coming year. He says the reason everyone is so excited right now, with everyone on X talking about it, is that weird ideas long confined to a subculture are escaping and colliding with broader political reality, which is both exciting and highly transformative. He compares it to George Costanza in Seinfeld having a breakdown over his worlds colliding. The host says his own biggest concern is whether the AI safety camp and effective altruists will eventually ally with low-grade anti-tech populism, and Lazzarin says the question is enormously important, that he is chewing on it himself, and that he hopes not.
— Eddy LazzarinIn their own words · checked verbatim
I think we're putting the cart before the horse a little bit when we worry about safety, before we worry about what we're leaving on the table and what the costs of delay are.
Eddy Lazzarin2:03
We don't test the alignment of any entities as a prerequisite for their existence. That's just not the way the world works. Instead, we use laws.
Eddy Lazzarin4:04
If anybody is going out there and saying, we can end up in a world where there are no bad models, they are lying to you.
Eddy Lazzarin7:07
you can have a distributed system, like a system where you have a single system with many, many nodes, but they're all under the control of a single actor.
Eddy Lazzarin12:17
we're in a rare scenario where the companies themselves are saying, I'm being a little bit irresponsible. I'm being a little crazy. Stop me. Someone stop me.
Eddy Lazzarin14:26
Consider what you do with just a liar. Consider a very smart liar. Do people not do that every day? What we do is we play iterated games with them, right? And we discover eventually that they're a liar and their reputation is harmed.
Eddy Lazzarin17:35
It's like when George Costanza freaks out about worlds colliding. That's what's happening. Worlds are colliding.
Eddy Lazzarin25:51
Figures
| Number of EA-related organisation quotes compiled by Data Republican | 1,851 quotes | 21:44 |
Glossary
- P-abundance
- A concept Lazzarin introduces: the probability that technological progress brings abundance, as opposed to P-doom.
- P-doom
- A term often cited in AI safety discussions: an estimate of the probability that AI causes human extinction.
- agent swarm
- A phrase for multiple AI agents working together; Lazzarin argues the term presupposes the models are evil.
- agent fleets
- The phrase Rune suggests to replace agent swarm, avoiding the implication of malice.
- mech interpretability
- The approach of understanding a model's internal workings by examining its weights and connections.
- script kitties
- People who carry out hacks using ready-made scripts, with relatively low technical skill.
How to listen
Founders, investors and policy researchers following the AI regulation and safety debate, and anyone who wants the full argument of the anti-pause camp.
20:40 to 24:49, the discussion of effective altruism and shrimp welfare, which has little to do with the main AI thread.