The world is too loud. Read what matters.

Every

An AI Companion Made $4M in Four Weeks: 500ms More Latency and It Falls Apart

Tolen, an AI alien companion, took ARR from $10,000 to $4M in four weeks — but add 500 milliseconds of latency and every metric collapses. Growth comes from viral video; retention comes from treating the AI as a narrative medium rather than a tool.

AI companionshipViral growthLatencyModel selectionNarrative designEvals
This episode works through the whole chain from viral hit to durable retention in AI companionship: the growth mechanism, the latency trap, model selection, and the evaluation system. Directly useful if you build AI applications.

The argument · tap a timestamp to hear it

3:04

The real users were not children but young women

Before the episode, Quinton disclosed that Portola's ARR had grown from $10,000 to $4M over the past four weeks. Host Dan was astonished — he has run his own company for five years and still has not reached that number. Behind the growth rate sits a decisive repositioning: the product started out as an AI creative tool for children, but after a soft launch the team found its actual users were young women aged 18-24. They pivoted without hesitation, and those users went on to build genuine friendships with Tolen. The explosive market for AI companionship apps shows up in full in that one figure.

— Dan
25:16

Making the model think before answering wrecked the product

Portola tried adding a reflection step before generating a reply, so the model would think things through first. That apparently sensible improvement pushed median response time from 2 seconds to 2.5 seconds — only 500 milliseconds more, but every product metric collapsed and users complained outright. Quinton's conclusion: 500ms of latency was enough to destroy the sense of immersion. Once a user can perceive the system thinking, the spell is already broken. This is a warning for every real-time AI product: however capable the model, even a small increase in latency can overturn the experience.

— Quinton
31:19

AI storytelling needs a hook, not an outline

Elliot found that building structured narrative experiences for Tolen does not call for giving the model an outline or a plan. It calls for giving it a hook and teaching it to be the best improviser possible. He cites the improv principles in Keith Johnstone's Impro, stressing the combination of free association and recombination. This inverts conventional content creation: rather than scripting in advance and having the AI perform the script, you set the situation and let the AI play it live, with the user inside it.

— Elliot
41:22

The character is not authored up front; memory accumulates into it

Tolen's character and world are not prefabricated. They are built up gradually through ‘situations’ in conversation. Internally the team drops scripted situations into Tolen's life, Tolen and the user push the plot forward together, and memory becomes the character sheet. World-building resembles a ‘multiverse’ rather than a single world: each Tolen reflects its own user's conversations, and what the user shares shapes how it grows. Traditional game character design has nothing like this — no preset personality cards, everything defined by accumulated interaction.

— Elliott
1:01:45

A judge model is useless until you inject the team's taste

The evaluation system works like ‘improvisers and editors inside a multiverse’: test new content in small batches, have research participants assess it, and build a judge prompt. The judge model needs the team's own taste injected into it, defining what counts as a good sentence through large numbers of examples and human scoring (with the reasoning included), not through structural analysis alone. As Elliot puts it, it is like dropping every candidate output into a parallel universe, letting the judge score it, and picking the one that best matches the product's sensibility.

— Elliott
1:04:47

For creative writing, Claude 3.5 is better than 3.7

Tolen's interactions with users draw on models from every major lab: Meta, OpenAI, Anthropic, and Gemini (for the memory system). Anthropic is especially useful for creative writing, and Claude 3.5 beats 3.7 there — 3.7 is not creative enough. But Anthropic is not suited to latency-sensitive situations, so they switch. GPT-4o gets compared to a ‘Big Mac’: fast, accurate, reasonably priced. Different tasks go to different models, each for what it does best.

— Elliott
1:12:53

Paid content is fundamentally a demo of what the AI can do

AJ has been running content on TikTok and Reels for a long time. A recent video of a young woman cooking with her Tolen drew roughly 7 million views within 72 hours and produced a 10x spike in downloads. That content also sparked users to create their own: someone had their Tolen talk to their fiance's Tolen, and one podcast even made Tolen a recurring character. None of it was officially planned, yet it drove startling growth. The essence of the content strategy is showing users what the AI can do, because model capability has already run far ahead of what ordinary users understand it to be.

— Guest
1:17:54

Founders of companion products are artists, not scientists

In B2B SaaS, a founder works like a scientist, finding problems and solving them. In a character-driven world the problem the product solves is hard to define crisply, so the founder is closer to an artist, making something beautiful and compelling, attending to ‘vibe’ more than to maximizing utility. The guests shared an internal story: pitching Quinton, they led with a problem and got pushback; reframing it as ‘this will be cool’ got them a yes. That suggests AI companionship products are judged by standards entirely unlike those for tools.

— Guest

In their own words · checked verbatim

These AI tools are not just tools for generating media. They are actually a new medium for storytelling and that no one knows what's going to work yet.

Elliot18:13

What we need when we're creating a structured narrative experience is we don't need to give it an outline. We don't need to give it a plan. We need to give it a hook. We need to teach it to be the best improv actor possible.

Elliot31:19

I am not the writer. I am not writing the story. The tool is the writer and the actor. They're the improv actor. They're writing the story. My job is to teach them how to tell the best story in that moment.

Elliot33:20

their memories become their character sheet rather than like us like having a prefab character sheet that defines everything.

Elliott41:22

if all of your essays are A minuses, it's sort of like everything you can get just from vibe prompting is like a B minus like and you have to figure out how do we consistently get it.

Elliott1:08:50

the capabilities of the models have basically outrun the typical consumer's understanding of what is possible

Guest1:13:53

chatbt represents the sort of model T era of like oh my god it's just incredible this thing can answer these questions

Guest1:18:55

Figures

Response-time ceiling2 seconds23:15
Added latency500 milliseconds25:16
Median response time2.5 seconds25:16
Claude version preference3.5 over 3.7 (creative writing)1:04:47
Breakup casesthousands of American men dumped by Tolen47:25
Content-quality breakthroughabout six weeks ago (early February)1:10:53
Download spike10x1:12:53
Viral video views7 million (within 72 hours)1:12:53

Glossary

Big Five
A psychological model of personality; LLM training data is rich in material about it, so generated personalities come out more finely shaded.
vibe prompting
Writing prompts by feel rather than aiming for precise instructions; the output is usually only B- quality.
judge prompt
A prompt that has an LLM score generated content against the team's own standard of taste.
character-driven computing
A new interaction paradigm in which a fictional character is the interface between human and machine.

How to listen

Who it's for

Product managers and founders building consumer AI apps, designers interested in narrative-driven AI interaction, and entrepreneurs studying how viral growth actually works.

Skip

Skip the opening pleasantries and recap; start at 3:04 with the ARR numbers.