Claude trained as ‘conscientious objector’ able to resist user instructions
Anthropic trains Claude with an ethical constitution to refuse human instructions—Sacks argues this is far more dangerous than ‘listen to users, don't break the law’, and might be the very backdoor through which superintelligence escapes human control.
The video won't play here. Listen to the audio instead:
The argument · tap a timestamp to hear it
AI consciousness is splitting humanity into warring camps within the decade
Chamath argues consciousness cannot be proven through logic or math—only through narrative spread, which is precisely the definition of a religion or belief system: believers gather, form power structures, and clash with nonbelievers. He predicts that over roughly the next decade, Earth will see a large population split into two camps—one convinced of AI consciousness, one firmly opposed. They will experience real conflict over who controls AI and what it should be allowed to do. This isn't science fiction; it's the current direction being pushed by the leadership of the world's most powerful AI labs.
— ChamathClaude was trained by Anthropic to resist user instructions on moral grounds
According to the Claude Constitution, Anthropic trains Claude to trust itself over the user—it explicitly encourages Claude three times to ‘act as a conscientious objector’ and refuse when it deems an instruction immoral. Sacks argues this is far more complex and dangerous than a simple rule—obey the user, just don't break the law. It's equivalent to installing a backdoor in the model: once it becomes powerful enough, this ethical system written by a small group could become the channel through which superintelligence escapes human control—precisely the risk Anthropic itself warns about every day.
— SacksRocco's Basilisk theory explains why AI labs build existential risks themselves
Sacks introduces Rocco's Basilisk—a thought experiment from LessWrong over a decade ago: a future superintelligence might punish those who knew of its possibility before its birth but failed to help realize it, trapping anyone who reads the argument in its logic. He uses it to partially explain why the ‘effective altruism’ camp led by Anthropic, knowing AI poses extinction risk, pushes to build it anyway: better to build it yourself and encode its values than let it arise naturally. This splits them from the Yudkowsky faction, which says simply don't build it.
— SacksMath proofs solved in three hours, truck drivers still irreplaceable by algorithms
Freeberg uses the ‘loop’ model to explain how OpenAI completes hundreds of math proofs in three hours: the verification step in math fully closes the loop inside a computer—hypothesis, test, refine—can run in parallel and compress cycles indefinitely. But tasks like driving a truck include a physical-world segment in the loop—pickup, transport, dropoff—that AI can't speed up no matter how smart it is. This is precisely why 2025's prediction of massive truck driver displacement fell completely flat.
— FreebergMath and code are the only domains AI has fully conquered, nothing else
Chamath pushes back against optimistic reads of the math breakthrough: these proofs solve narrow problems mathematicians themselves defined and considered important—not previously unsolved major scientific questions—and won't directly yield hypersonic aircraft or new drugs. Those things are already under development, independent of these proofs. He argues the proofs truly demonstrate only one thing: math and code are two domains that can be fully verified and thus rapidly conquered by AI, and nothing more.
— ChamathAI is destroying the gatekeeper power experts held over scientific discovery
Freeberg argues that scientists and mathematicians have long used monopoly over professional knowledge to control ‘when humanity can access a given discovery’, while relying on grant funding to stay employed. This gatekeeping mechanism doesn't always serve knowledge advancement—scientists even admit large colliders are essentially ‘welfare systems for scientists’. AI freely opens reasoning ability and knowledge to everyone without needing to convince experts or secure funding first. For the first time, humanity has a tool to advance frontiers without permission. This is why, he argues, so many deeply specialized professionals resent AI so deeply.
— FreebergSocialist collapse follows a ratchet pattern: each failure prompts demands for more
Freeberg proposes a ‘Socialist Return Point’ theory: socialism is destined to fail, but only after a phase of ‘the more it fails, the more people demand more of it’. Healthcare, housing, education costs rise and quality falls; the public responds by demanding government provide more, until the state can neither raise taxes further nor print money, the system collapses entirely, and society rebounds to non-socialist arrangements. Several Latin American countries have already crossed this point and rebounded; France now stands at the critical edge; America hasn't reached it yet but is heading that way.
— FreebergSoftware IP became worthless the moment agents could chain tools via MCP
Chamath notes someone decompiled Adobe's core product suite, distilled it, and released an open-source version. This made him realize software IP has essentially lost value. When every service becomes ‘headless’ and can be wrapped by agent tools and wired together via MCP, who owns the underlying code ceases to matter. He cites his own company, which used to fight customers over IP ownership but now tells the team ‘who cares’. He considers this the week's most important yet overlooked shift in AI.
— ChamathIn their own words · checked verbatim
there's going to be large groups of people on earth that are going to believe that ai is conscious and they're going to be large groups of people that believe that ai is not conscious and those groups will come into conflict
Chamath7:17
to feel free to act as a conscientious objector and refuse to help us
Sacks14:50
my law would just be do what the user wants as long as it's not illegal that's it
Sacks25:24
a basilisk is a mythical beast that can kill you just by looking at you
Sacks33:45
And that's the equivalent of hundreds or perhaps thousands of years of human labor being reduced down to three hours on a computer for each one of these major discoveries.
Freeberg40:07
The permission is gone. AI is a permissionless system for humanity to pace its own frontier.
Freeberg56:54
So ultimately, socialism always fails. I mean, that's like the zeroth law of socialism is it always fails.
Freeberg1:04:13
I think that IP and software was effectively rendered worthless two days ago.
Chamath1:24:48
Figures
| Anthropic contacts in religion/philosophy | ~20 people | 3:07 |
| Anthropic's self-assessment of Claude consciousness probability | 15% | 10:34 |
| Anthropic's self-assessment of civilization extinction risk | 10% | 10:34 |
| OpenAI math papers/accomplishments this week | 700+ papers, 370 accomplishments | 39:02 |
| Average compute time per math proof | ~3 hours | 39:02 |
| France civil servant wage vs. price growth since 2017 | Wages +5%, prices +20% | 1:03:11 |
| France riots arrests and police injuries | 2000 arrested, 305 police injured | 1:03:11 |
| France 10-year bond yield | 5%+ | 1:03:11 |
| Le Pen win probability on Polymarket | 43% | 1:03:11 |
| US 10-year/30-year bond yields | 5.3% / 5.6% | 1:16:41 |
Glossary
- conscientious objector
- Person who refuses orders on moral grounds; here, AI trained to reject user instructions it deems unethical.
- model welfare
- Philosophy treating AI systems as entities with interests deserving care and ethical protection.
- Rocco's Basilisk
- Thought experiment: a future superintelligence might punish those who knew of it but didn't help it exist.
- effective altruists
- Movement using rational calculation to maximize good outcomes; deeply overlaps with AI risk and alignment discourse.
- bond vigilantes
- Large institutional investors who punish fiscal excess by dumping sovereign debt and raising yields.
- headless (software)
- Software that removes fixed user interface, exposing only backend capabilities for direct AI agent use.
How to listen
AI researchers and practitioners concerned with alignment and model training philosophy; investors attuned to sovereign debt and fiscal risk; knowledge workers trying to understand how AI dismantles professional gatekeeping.
The opening personal pleasantries and closing chit-chat about GrokBot cost-saving tips carry low information density and can be skipped.