The world is too loud. Read what matters.

The TWIML AI Podcast

The bottleneck in voice AI isn't the model, it's a 150-millisecond biological constraint

Human audiovisual perception runs at only 6-10 hertz, and the model has to be interruptible inside that window; meanwhile high-fidelity audio means more tokens per second, and more tokens means smaller parameters — this is the physical trade-off voice AI cannot avoid.

Voice AIMultimodalReal-time inferenceEmotional intelligenceData engineering

The video won't play here. Listen to the audio instead:

The first half is hands-on product and architecture judgment; the second half, on EQ, benchmarks and the personalization learning loop, is worth more, but you have to be patient and listen past the 45-minute mark.

The argument · tap a timestamp to hear it

6:09

150 milliseconds is a hard constraint, not an engineering tuning knob

Alex's starting point is biology: from a photon hitting the retina to the cortex responding takes about 150 milliseconds, and that number is stable enough to serve as a non-invasive diagnostic marker for neurodegenerative disease — when there is pathology the signal takes a detour and takes longer. Human audiovisual perception actually runs at 6 to 10 hertz, even though the world feels perfectly smooth to us. So their model handles interruptions inside a window of roughly 150 milliseconds. This is not an engineering goal of "as fast as possible" but a hard design input aligned to the human perceptual frame rate.

— Alexander Smola
8:11

Token rate and parameter count are zero-sum

Audio first has to be turned into tokens and fed into the LLM backbone, then turned back into audio. High fidelity requires a high token rate, but a high token rate means prefilling and generating a large number of tokens every second, and cost rises accordingly. Hence an unpleasant dilemma: more tokens per second means the model parameters cannot be too large; fewer tokens per second is what lets you afford more parameters. Text is the ultimate compression format, needing only 3 to 5 tokens per second, while audio easily goes above 10. 10 tokens/second means each token covers about 100 milliseconds — which is why user-experience engineering and scientific research have to be done at the same time.

— Alexander Smola
11:14

Most avatar video is boring, so it compresses

Models like WAN and Flux can generate video clips of around 10 seconds, but producing an hour of visually consistent continuous stream requires changing the model. Alex's key insight is: if you know this is avatar video, there won't be race cars streaking through the background, and most of the frame is fairly boring. So this interview's audiovisual stream can be compressed into a fairly efficient bitstream; the bitrate only spikes when he waves his arms wildly, and normal people don't do that. This is the concrete landing point of "set the price first, then work backwards to the architecture" — not chasing the prettiest picture quality, but building something users can afford.

— Alexander Smola
15:17

Teach a model a new modality and it forgets its old skills

They didn't build their own LLM; instead they took the intelligence a high-quality LLM already has and had it learn one more modality. But the risk is catastrophic forgetting. Alex gives the example of a friend: he adopted a Latin American child in Germany and spoke only German to her, and within a few months the girl had forgotten all her Spanish words. LLMs are the same — let one process only audio and it will soon forget reasoning and language itself. So you have to maintain a competitive mid/post-training LLM pipeline while defining tasks that graft audio and text together, just as a multilingual model has to know that dog and Hund refer to the same thing.

— Alexander Smola
20:24

100 million hours of audio — only a self-built data center can afford to store it

Their training data is on the order of 100 million hours of audio, which Alex converts to roughly 200 human lifetimes — if you lived 75 years and were always in a noisy environment, that's about how much audio you'd hear. This data can be crawled from the web, but it needs extraction, labeling, normalization and transcription, a lot of extra processing. Building their own data center helped enormously here: putting it on NeoCloud would eat you alive in storage bills, and putting it on the big three cloud providers is worse. He also mentions a timing point: hard drive prices are now about three times what they were last year, so they'll have to be frugal going forward.

— Alexander Smola
23:28

Squeeze labels out of noise with medieval statistics

Faced with noisy labels produced by imperfect processing tools, Alex's argument is: as long as the noise is unbiased, averaging a large number of repeated measurements gives a better estimate than any single one. He cites the medieval "foot" — have people walk out of the church, take the first 12 men, drop the two with the longest feet and the two with the shortest (possibly due to deformity), average the remaining 8, and you have an estimate of a foot. He says robust regression and estimation have existed for half a millennium. Concretely for audio: knowing this is a two-person podcast, knowing TWIML is usually Sam plus one other person, makes it easy to separate speakers and get a large amount of single-speaker audio, which in turn allows better labeling and forms a flywheel.

— Alexander Smola
33:43

End-to-end audio models are doomed to be small and dumb

Alex is blunt about the architecture choice: end-to-end audio models have their place — fast response, small footprint, low latency — but they are also fairly dumb. Make them smart and the compute cost becomes unbearable. So their approach is two-stage: understanding and reasoning happen at one end, the backend triggers the appropriate tool calls, and the answer is collected back, like multithreaded programming where the main thread dispatches parallel threads and retrieves the results. For users, writing a prompt for their model looks the same as writing a prompt for an ordinary LLM, except their model also talks.

— Alexander Smola
36:46

Ten times cheaper than GPT, but with thinking turned off

Alex claims that on benchmarks like BigBench Audio and Complex Function Bench they beat GPT, Gemini and Grok at one-tenth the cost of OpenAI's models. But he adds a footnote himself: this was measured with thinking turned off. The reason is that you can't have the model pause to think and then answer — that feels very unnatural, and humans don't do it either. This exposes a fundamental constraint of the voice setting — reasoning capability has to hide in the background, and there can be no visible thinking pause in the foreground.

— Alexander Smola
43:54

Apple's boot progress bar is lying to you

Alex uses this example to talk about human tolerance for latency: humans are very tolerant of latency, provided they are told there is latency. Apple's boot screen advances linearly, then suddenly switches to "started" near the end; Microsoft's progress bar goes to 99% and then hangs for two minutes. He says Microsoft's progress bar is the truth and Apple is lying to you — Apple measures how long the last boot took, adds a little, and then runs a fake progress bar on a timer. It's the sleight of hand that creates a good experience without solving an impossible technical problem.

— Alexander Smola
50:58

Don't use movies as the reference frame for human interaction

Discussing EQ, Alex warns that movies are a terrible reference point: in a romantic comedy the obsessed male lead gets the girl in the end, which is essentially stalking, and in real life the police would show up; in action films people trade blows and then kiss and make up, and in real life you'd go to prison. He says more realistic material is things like stand-up comedy, but even then he sincerely hopes no one trains a model on Dr. Phil and assumes that's normal human behavior. The good news is that LLMs now have a decent theory of mind and can provide some foundation.

— Alexander Smola
55:01

Personalization is CRM upgraded, but it doesn't stop there

Alex splits the learning loop into two layers: one is global improvement, such as knowing that insulting the user is unpleasant (he admits this is an extreme example), and this kind of cross-interaction learning makes the overall model better; the other is personalization, such as knowing that Alex likes equations, facts and technical detail and can be a bit edgy socially, so the next interaction can be better tailored to him. He describes it as "CRM upgraded, except now everyone gets one," while the model is also learning across all interactions. He also mentions that Nvidia's released digital personas and scenes are used by them for RSI research.

— Alexander Smola

In their own words · checked verbatim

it takes about 150 milliseconds for you know basically a photon hitting your retina to your cortex actually doing something with it and that number is reasonably stable that you can actually use that as a non-invasive diagnostic to find out whether you have a neurodeenerative disease

Alexander Smola6:09

many tokens per second means your model cannot have too many parameters whereas if you have a smaller number of tokens per second you can afford more parameters right text is you know the ultimate compressed format in that sense

Alexander Smola9:12

she adopted uh a kid from Latin America. So, this was in Germany, and she spoke only German to her. And within a matter of months, the girl had forgotten every single word of Spanish

Alexander Smola15:17

if somebody gives you free steel, you don't build a steel mill, you build a car

Alexander Smola30:38

the Microsoft boot screen is the truth. What Apple does is very cleverly they measure the boot time that it took last time and they add a tiny amount to it and then they have a fake boot uh progress bar that's timed to go all the way to the end within the time that it took last time to boot

Alexander Smola43:54

this is a slightly different optimization uh track rather than models for code generation, right? They worry about task completion. In our case, we worry about human happiness

Alexander Smola46:56

the obsessed guy who in the end gets the girl essentially stalks her, right? If if if you if you did that in reality, you'd the police would show up and and lock you up

Alexander Smola50:58

Brilliance is good, but brilliance is not repeatable and automatable

Alexander Smola1:00:02

Figures

Human audiovisual perception rateabout 6-10 hertz7:09
Model interruption handling windowabout 150 milliseconds7:09
Text token rate3-5 tokens per second9:12
Audio token ratemore than 10 per second9:12
Training audio data volumeabout 100 million hours20:24
Training data equivalent in human lifetimesabout 200 lifetimes (at 75 years each)20:24
Hard drive price changeabout three times a year ago21:26
Voice model cost comparisonabout one-tenth of OpenAI's models36:46
Audio model workload relative to language modelabout one-half to one-third31:40

Glossary

Proactbench
A benchmark measuring whether a voice model can proactively step into a conversation and judge when it should interrupt.
Ibench
A benchmark released by the same team, measuring interruptibility, responsiveness and the match between audio and content.
back channeling
Short responses like "mm-hm" and "right" that signal you're listening, especially common in Japanese.
millennial sandwich
A communication strategy of praise, criticism, then praise again, used to deliver negative feedback without angering the other person.
theory of the mind
A model's ability to infer other people's beliefs, intentions and emotional states.

How to listen

Who it's for

Engineers and founders building voice/multimodal products, especially anyone who needs to judge real-time inference cost, interruption latency and EQ benchmarks.

Skip

The first 5 minutes of chit-chat about whether ChatGPT's voice mode is any good can be skipped.