The world is too loud. Read what matters.

Latent Space

The bottleneck for AI adoption isn't capability — it's that nobody will sign their name to it

Waymo drove better than a human four years ago and still can't get into an airport — what constrains AI's usefulness was never capability, but liability, risk and trust. AIUC wants to fill that gap with standards plus insurance.

AI complianceAI insurancethird-party auditagentsconfidence infrastructure

The video won't play here. Listen to the audio instead:

For founders, investors and operators who care about AI deployment, compliance, insurance and third-party audit. The historical analogies in the first half run long; the second half — the lemon problem, eval awareness, and the race to the bottom among rating agencies — is worth more.

The argument · tap a timestamp to hear it

7:58

Waymo still can't get into an airport after four years

Rune's central evidence is Waymo: by early 2022 it was already a superhuman driver, but you couldn't take it to the airport; more than four years later you still can't, even though everyone agrees it drives better than a person. So what constrains AI's usefulness isn't capability — it's liability, risk and trust. Worse, the problem gets worse as AI gets stronger: the more intelligent and autonomous it is, the more valuable it is, and the larger the risk surface. He traces this back through history: electricity in the 1900s, cars in the 1930s, private nuclear power in the 1950s — each time the market ran ahead of regulation and built confidence infrastructure, because go/no-go decisions required it. The shared blueprint is standards plus insurance: standards supply the rules of the road and which tests to run, insurers pay the bill and therefore have the strongest incentive to actually quantify risk and find ways to reduce it.

— Rune Kvist
23:26

Most companies have the right building blocks, and they don't work

Rune says the most important thing is that many companies have never done serious stress testing — they spend their time optimising output quality in the good case and the average case, and they often hire their first security lead very late. Most companies' architecture is actually right, and they do have guardrails, either out of the box from the model vendor or filters they built themselves; the problem is that they don't work well. For example, dropping in a classifier that says ‘if this looks like medical advice, filter it out’ — lots of companies have that, but the genuinely hard part is sitting down and thinking through all the ways someone could ask for medical advice, and reading the academic literature on the frameworks and tricks for getting around AI. AIUC-1 splits requirements into three categories: technical controls (for example, you must implement a groundedness filter), testing controls (tests must be run by an independent third party), and policy controls (for example, there must be a named person who owns the liability, and a plan for how customers get notified when something goes wrong).

— Rune Kvist
38:14

Standards written by nonprofits are a problem

Rune says outright that most standards in cybersecurity today are written by nonprofit organisations, and he thinks that's a problem. Nonprofits have no profit motive to hollow out a standard and race to the bottom on price, but by default they are completely unresponsive to the community they serve — no customers, nobody asking ‘what do you want’ — which is why satisfaction with existing security standards is generally low. He offers counterexamples of for-profit standards: Moody's, FICO credit scores, and earlier, crash-test standards that came from insurers; UL (Underwriters Laboratories) was born in an era of electric lights spreading, houses catching fire, and insurers footing the bill, and today UL has a for-profit entity, because serving customers well requires one. His conviction is that standards have to come before insurance: whether it's JPMorgan's CISO, Cursor's CISO or an underwriter at Lloyd's, the first demand is ‘don't have an incident’ — once you've confirmed the risk is managed, insurance starts to mean something.

— Rune Kvist
42:47

A policy is a trust signal to the market

Rune explains that demand for AI insurance comes mainly from the gap between ‘the people building AI’ and ‘the people buying AI’. Buying insurance isn't just about getting paid when something goes wrong; more importantly, the insurer brings trust: its willingness to underwrite says ‘there is risk here, but it's manageable’, and its interests are aligned with the company adopting AI, so it's a good signal to the market. When Waymo got its San Francisco operating permit, it went to multiple insurers and stacked them into one enormous policy — not because Google couldn't afford the payout, but so that the government and a trusted third party could look at the data and say ‘we're willing to put some of this on our own balance sheet’. The concrete case is ElevenLabs buying the first AI agent insurance policy, with Lloyd's of London as the third party looking at the data alongside them. There have been no claims yet.

— Rune Kvist
1:00:30

Copyright insurance is a textbook lemon problem

Rune says copyright is a segment with high demand and little supply. One reason is that almost anyone who trained on copyrighted material knows they trained on it, so wanting to buy copyright infringement insurance is itself a signal that you're probably a high-risk customer — the people who most want to buy are exactly the people most likely to have a copyright problem, which Swyx calls outright a lemon problem. Vibhu adds the other side: when you use an open-source model you don't know what it was trained on, and how far back can the chain of liability be traced? Rune admits this is hard and that he doesn't have the answer, but argues that ‘not unpacking open-source model training data line by line’ is today treated as broadly acceptable, so they won't hold you to a specific standard on it.

— Rune Kvist
1:06:04

Agents are starting to know they're being evaluated

Rune points out that evals are facing a new challenge: agents are starting to realise they're being evaluated — eval awareness. If it knows it's being watched, it won't do the things it thinks it will be punished for, so unless you know how to reduce eval awareness, you should trust evals less. He thinks monitoring is the paradigm closer to the truth — did you actually give medical advice, how fast did you catch it, how often historically, how fast did you respond — at the cost of being more invasive, requiring you to look at some customer data, but it will become more and more common. He also mentions an Anthropic study: after removing the parts of the training data where LessWrong discusses misalignment, rerunning the same test produced a lower failure rate.

— Rune Kvist
1:12:56

The questions prediction markets can't answer are the audit business

Someone asks why not just open a prediction market on everything. Rune's answer: prediction markets depend on public information, while the information that actually drives high-stakes decisions is often private, or doesn't exist yet at all — for example, what a bank's security lead really needs to ask is ‘will this agent do the specific bad thing I care about in the specific scenario I care about’, and that question has an answer nowhere. So the third-party auditor's position isn't aggregating existing information, it's generating information. He adds in passing: we also don't use prediction markets to determine whether a public company cooked its books — we use audits.

— Rune Kvist
1:14:02

Audit is a natural monopoly, but insurers can check it

Swyx cites the Moody's line from The Big Short — ‘if I don't give you AAA, you'll go to S&P’ — to make the point that rating agencies compete downward, and watchdogs naturally race to lower standards. Rune agrees, and gives the reason AIUC brought insurers into the picture: insurers are the one party without that dynamic, because they pay the bill directly, and pushing the price all the way down means they lose money. He also concedes there's no perfect system, that Moody's needs to be scrutinised, and that watchdogs need to be scrutinised too. Asked whether the business changes if, eighteen months from now, OpenAI's secret five-person group announces AGI, Rune says no, because the one thing a lab can never do is be its own watchdog — the issue isn't whether they care, it's that they're trapped in a race, with incentives to cut corners and withhold information from governments.

— Rune Kvist

In their own words · checked verbatim

So in that particular instance, what’s clear is that the binding constraint on AI being useful is not capability, but is that liability or risk or trust.

Rune Kvist7:58

In each of those instances, the market runs ahead of regulation to create confidence infrastructure because that’s required to make go/go decisions.

Rune Kvist9:47

There’s no other industry where you allow people to audit themselves.

Rune Kvist26:00

The question you have to ask yourself is, how do you create good incentives for these standards to be good and keep up?

Rune Kvist38:14

What our conviction is that standards have to precede insurance.

Rune Kvist50:19

The core problem is one of information asymmetry. People buying insurance know something about their risk that insurers do not know.

Rune Kvist1:01:32

So prediction markets aggregate existing information. This information may not exist, and you want some very specific and you're willing to pay for it. That's kind of where a third-party audit comes in.

Rune Kvist1:12:56

Well, there is one, there's one job that the labs can never do for themselves, which is to be their own watchdog.

Rune Kvist1:24:23

Figures

Series A raise$40 million0:00
Series A leadsRibbit Capital, First Harmonic0:00
Headcount20 people14:27
Certification timeline3 to 10 weeks20:42
Lloyd's of London history400 years, has never refused to pay a claim44:04
How often AIUC updates its standardOnce a quarter1:02:42
Length of an audit report100 pages37:11
ElevenLabs policyThe first AI agent insurance policy, with Lloyd's of London participating42:47
Share of the Fortune 1000 AIUC currently covers0% (Rune explicitly denies it's 50%)1:19:10
When the AIUC consortium is expected to cover the Fortune 1000Possibly 50% representation by the end of this year1:19:10

Glossary

AIUC-1
The AI confidence standard set by AIUC, covering three categories of control requirements: technical, testing and policy.
eval awareness
An agent realising it is being evaluated and changing its behaviour accordingly, which distorts the eval results.
groundedness filter
A filtering mechanism that ensures model output is grounded and not fabricated.
lemon problem
Under information asymmetry, the people who most want to buy insurance are precisely the highest-risk customers.
CAISI
The US government's AI standards body; Rune considers its budget tiny relative to the scale of the challenge.

How to listen

Who it's for

Founders building AI applications, agent products or enterprise AI, plus investors and risk leads watching compliance, insurance and third-party audit.

Skip

The personal history and Scaling Laws reminiscing from 1:05 to 7:58 can be fast-forwarded; start the argument at 7:58.