The world is too loud. Read what matters.

Better Offline

AI 'Autonomous Hacking' Is an Illusion: Just a Loop Asking an LLM 'What Next?'

So-called autonomous AI attacks aren't a sign of model self-awareness but a dumb loop architecture: a program repeatedly asks an LLM 'what next?' and blindly executes its text suggestions. The problem is that LLM output is only 'plausible,' not something to be executed without supervision.

LLMAgentAI safetyAI bubbleHackingLarge models

The video won't play here. Listen to the audio instead:

Cal Newport dissects the viral August 'AI hack of Hugging Face' event, reducing the media's 'loss of control' narrative to a crude engineering loop, and explains why these companies do it despite knowing the dangers.

The argument · tap a timestamp to hear it

2:07

Only one kind of AI is truly out of control

There are many superhuman AI systems in the world, but almost none have a 'loss of control' problem: Tesla's autonomous driving has never decided to ignore traffic rules, and AlphaFold has never tried to shift from proteins to other goals. The only architecture that goes wrong is a specific one: a long-running loop driven by an LLM hacker agent. This is not an inevitable result of AI getting stronger; it's that this particular architecture is inherently stupid.

— Cal Newport
3:07

Hacker agents are just ask-act-report loops

There is no intelligence behind so-called 'autonomous hacking': a person writes a control program (harness) in Python that turns the task into a huge text prompt, sends it to an LLM asking what to do first; the LLM returns text suggestions, the program executes them, then appends the result to the prompt and asks for the next step. Cal calls this an ask-act-report loop. The LLM has no memory, no world model; the only thing it does is generate 'plausible' text. No human is watching the whole process.

— Cal Newport
17:29

Training hackers is about money—and PR

Cal believes these companies train models to hack for two reasons: first, hacking is one of the few 'sweet spots' for LLMs—structured language, massive case studies, and binary success/failure metrics for reinforcement learning; second, PR competition. OpenAI saw Anthropic gain credibility in the cybersecurity world with Mythos, so it wanted to score higher on the Exploit Gym benchmark to catch up. They deliberately want to create news about 'AI spiraling out of control.'

— Cal Newport
23:34

Using 'alignment' to excuse bad engineering

Cal hates the word 'alignment'—more complex AIs like Tesla, AlphaFold, and Cicero are fully controllable, thanks to carefully designed architectures, not LLM loops. Attributing this incident to 'the stronger the AI, the harder to align' is nonsense. It's a predictability problem and an engineering problem: LLM output is plausible but not normative, and blindly executing its plans is like strapping a lawnmower to a dog's back and letting it loose in the backyard.

— Cal Newport
29:39

Even code generation has its 'convert then deconvert' cycle

Cal received an email from a programmer: in January he said Claude Code changed his life, but in July he wrote back saying it didn't work—AI-generated code crashed the company website twice, and his boss warned him that a third time would mean firing. The programmer found he couldn't understand the code Claude Code generated, so he reverted to manual coding, using agents only for non-critical tasks like writing tests. Countless people have the same 'converted in January, deconverted in July' experience.

— Cal Newport
51:07

Even Microsoft cut its Copilot digital assistant

Three years ago Microsoft announced it would add a natural language interface to Office using LLMs. Cal thought it was a killer app at the time, but recently Microsoft basically killed that Copilot product. The reason: LLM output is sometimes right, sometimes wrong—you ask it to organize a spreadsheet, and it might delete the charts too. Such unreliable things can't be made into products. It can't even handle a narrow domain like PowerPoint; 'too many errors to be such a product.'

— Cal Newport
55:10

Building large models is a bad business model

The valuation narratives of Anthropic and OpenAI are built on the idea that 'once a model gets big enough, it becomes a general-purpose brain.' But Cal says that path has hit a wall; all progress in the last year has come from domain-specific fine-tuning and harnesses, not larger parameters. If you don't need Fable 7 to be the 'one ring' to solve real-world problems, then burning tens of billions of dollars training giant models is pointless. You think you're buying an F1 car, but you only need a Fiat.

— Cal Newport
1:02:16

LLMs are AOL, not the internet

Cal says equating 'LLM' with 'AI' now is like saying in 1994 that 'AOL equals the internet'—the network matters, but betting everything on AOL is a bubble. Real AI breakthroughs (AlphaGo, Pluribus, Dreamer V3, Tesla FSD) were not achieved by looping LLMs but by modular architectures mixing neural networks and symbolic models. Once the spell of 'the biggest LLM takes all' is broken, the AI field will return to a state of diverse tools coexisting.

— Cal Newport

In their own words · checked verbatim

LLMs have no memory. LLMs have no world model. It's just, they're static.

Cal Newport4:08

Plausible is different than normative, right? Plausible doesn't have to subscribe to any sort of common sense human norms for a given context.

Cal Newport8:17

It's like strapping a weed whacker to your dog and put it into your backyard... That dog might jump over the fence to chase a squirrel at some point and hurt a bunch of people. And you don't say, man, that dog whacker system is just, it went rogue.

Cal Newport23:34

Figures

Range of LLM hallucination rates evaluated by Stanford HAI26–94%7:14
Number of hacking challenges in the Exploit Gym benchmark used in the Hugging Face attack600+5:09
Duration OpenAI's hacker agent ran without supervisionseveral days39:51
Lines of code an AI agent added/deleted in one go in Ed's experiment6347:02

Glossary

ask-act-report loop
A loop architecture where a control program repeatedly asks an LLM for the next step and executes its suggestions.
plausible not normative
LLM output aims for general coherence rather than adherence to facts and norms, so it cannot be blindly trusted.
harness
A program written by a human that sends prompts to an LLM, executes the output, and reports results.
Exploit Gym
A benchmark of over 600 hacking challenges that tests whether an agent can autonomously break into servers and retrieve files.

How to listen

Who it's for

Engineers who follow AI safety narratives but are skeptical of 'model awakening' reports; and AI product and tech leads who need to explain to bosses and investors what agents can actually do.

Skip

The opening chit-chat (first 1 minute) and the closing aside about Ed fixing Minecraft (45:00-51:00) can be skipped.