The bottleneck in voice AI isn't the model, it's a 150-millisecond biological constraint
Human audiovisual perception runs at only 6-10 hertz, and the model has to be interruptible inside that window; meanwhile high-fidelity audio means more tokens per second, and more tokens means smaller parameters — this is the physical trade-off voice AI cannot avoid.
The video won't play here. Listen to the audio instead:
The argument · tap a timestamp to hear it
150 milliseconds is a hard constraint, not an engineering tuning knob
Alex's starting point is biology: from a photon hitting the retina to the cortex responding takes about 150 milliseconds, and that number is stable enough to serve as a non-invasive diagnostic marker for neurodegenerative disease — when there is pathology the signal takes a detour and takes longer. Human audiovisual perception actually runs at 6 to 10 hertz, even though the world feels perfectly smooth to us. So their model handles interruptions inside a window of roughly 150 milliseconds. This is not an engineering goal of "as fast as possible" but a hard design input aligned to the human perceptual frame rate.
— Alexander SmolaToken rate and parameter count are zero-sum
Audio first has to be turned into tokens and fed into the LLM backbone, then turned back into audio. High fidelity requires a high token rate, but a high token rate means prefilling and generating a large number of tokens every second, and cost rises accordingly. Hence an unpleasant dilemma: more tokens per second means the model parameters cannot be too large; fewer tokens per second is what lets you afford more parameters. Text is the ultimate compression format, needing only 3 to 5 tokens per second, while audio easily goes above 10. 10 tokens/second means each token covers about 100 milliseconds — which is why user-experience engineering and scientific research have to be done at the same time.
— Alexander SmolaMost avatar video is boring, so it compresses
Models like WAN and Flux can generate video clips of around 10 seconds, but producing an hour of visually consistent continuous stream requires changing the model. Alex's key insight is: if you know this is avatar video, there won't be race cars streaking through the background, and most of the frame is fairly boring. So this interview's audiovisual stream can be compressed into a fairly efficient bitstream; the bitrate only spikes when he waves his arms wildly, and normal people don't do that. This is the concrete landing point of "set the price first, then work backwards to the architecture" — not chasing the prettiest picture quality, but building something users can afford.
— Alexander SmolaTeach a model a new modality and it forgets its old skills
They didn't build their own LLM; instead they took the intelligence a high-quality LLM already has and had it learn one more modality. But the risk is catastrophic forgetting. Alex gives the example of a friend: he adopted a Latin American child in Germany and spoke only German to her, and within a few months the girl had forgotten all her Spanish words. LLMs are the same — let one process only audio and it will soon forget reasoning and language itself. So you have to maintain a competitive mid/post-training LLM pipeline while defining tasks that graft audio and text together, just as a multilingual model has to know that dog and Hund refer to the same thing.
— Alexander Smola100 million hours of audio — only a self-built data center can afford to store it
Their training data is on the order of 100 million hours of audio, which Alex converts to roughly 200 human lifetimes — if you lived 75 years and were always in a noisy environment, that's about how much audio you'd hear. This data can be crawled from the web, but it needs extraction, labeling, normalization and transcription, a lot of extra processing. Building their own data center helped enormously here: putting it on NeoCloud would eat you alive in storage bills, and putting it on the big three cloud providers is worse. He also mentions a timing point: hard drive prices are now about three times what they were last year, so they'll have to be frugal going forward.
— Alexander SmolaSqueeze labels out of noise with medieval statistics
Faced with noisy labels produced by imperfect processing tools, Alex's argument is: as long as the noise is unbiased, averaging a large number of repeated measurements gives a better estimate than any single one. He cites the medieval "foot" — have people walk out of the church, take the first 12 men, drop the two with the longest feet and the two with the shortest (possibly due to deformity), average the remaining 8, and you have an estimate of a foot. He says robust regression and estimation have existed for half a millennium. Concretely for audio: knowing this is a two-person podcast, knowing TWIML is usually Sam plus one other person, makes it easy to separate speakers and get a large amount of single-speaker audio, which in turn allows better labeling and forms a flywheel.
— Alexander SmolaEnd-to-end audio models are doomed to be small and dumb
Alex is blunt about the architecture choice: end-to-end audio models have their place — fast response, small footprint, low latency — but they are also fairly dumb. Make them smart and the compute cost becomes unbearable. So their approach is two-stage: understanding and reasoning happen at one end, the backend triggers the appropriate tool calls, and the answer is collected back, like multithreaded programming where the main thread dispatches parallel threads and retrieves the results. For users, writing a prompt for their model looks the same as writing a prompt for an ordinary LLM, except their model also talks.
— Alexander SmolaTen times cheaper than GPT, but with thinking turned off
Alex claims that on benchmarks like BigBench Audio and Complex Function Bench they beat GPT, Gemini and Grok at one-tenth the cost of OpenAI's models. But he adds a footnote himself: this was measured with thinking turned off. The reason is that you can't have the model pause to think and then answer — that feels very unnatural, and humans don't do it either. This exposes a fundamental constraint of the voice setting — reasoning capability has to hide in the background, and there can be no visible thinking pause in the foreground.
— Alexander SmolaApple's boot progress bar is lying to you
Alex uses this example to talk about human tolerance for latency: humans are very tolerant of latency, provided they are told there is latency. Apple's boot screen advances linearly, then suddenly switches to "started" near the end; Microsoft's progress bar goes to 99% and then hangs for two minutes. He says Microsoft's progress bar is the truth and Apple is lying to you — Apple measures how long the last boot took, adds a little, and then runs a fake progress bar on a timer. It's the sleight of hand that creates a good experience without solving an impossible technical problem.
— Alexander SmolaDon't use movies as the reference frame for human interaction
Discussing EQ, Alex warns that movies are a terrible reference point: in a romantic comedy the obsessed male lead gets the girl in the end, which is essentially stalking, and in real life the police would show up; in action films people trade blows and then kiss and make up, and in real life you'd go to prison. He says more realistic material is things like stand-up comedy, but even then he sincerely hopes no one trains a model on Dr. Phil and assumes that's normal human behavior. The good news is that LLMs now have a decent theory of mind and can provide some foundation.
— Alexander SmolaPersonalization is CRM upgraded, but it doesn't stop there
Alex splits the learning loop into two layers: one is global improvement, such as knowing that insulting the user is unpleasant (he admits this is an extreme example), and this kind of cross-interaction learning makes the overall model better; the other is personalization, such as knowing that Alex likes equations, facts and technical detail and can be a bit edgy socially, so the next interaction can be better tailored to him. He describes it as "CRM upgraded, except now everyone gets one," while the model is also learning across all interactions. He also mentions that Nvidia's released digital personas and scenes are used by them for RSI research.
— Alexander SmolaIn their own words · checked verbatim
it takes about 150 milliseconds for you know basically a photon hitting your retina to your cortex actually doing something with it and that number is reasonably stable that you can actually use that as a non-invasive diagnostic to find out whether you have a neurodeenerative disease
Alexander Smola6:09
many tokens per second means your model cannot have too many parameters whereas if you have a smaller number of tokens per second you can afford more parameters right text is you know the ultimate compressed format in that sense
Alexander Smola9:12
she adopted uh a kid from Latin America. So, this was in Germany, and she spoke only German to her. And within a matter of months, the girl had forgotten every single word of Spanish
Alexander Smola15:17
if somebody gives you free steel, you don't build a steel mill, you build a car
Alexander Smola30:38
the Microsoft boot screen is the truth. What Apple does is very cleverly they measure the boot time that it took last time and they add a tiny amount to it and then they have a fake boot uh progress bar that's timed to go all the way to the end within the time that it took last time to boot
Alexander Smola43:54
this is a slightly different optimization uh track rather than models for code generation, right? They worry about task completion. In our case, we worry about human happiness
Alexander Smola46:56
the obsessed guy who in the end gets the girl essentially stalks her, right? If if if you if you did that in reality, you'd the police would show up and and lock you up
Alexander Smola50:58
Brilliance is good, but brilliance is not repeatable and automatable
Alexander Smola1:00:02
Figures
| Human audiovisual perception rate | about 6-10 hertz | 7:09 |
| Model interruption handling window | about 150 milliseconds | 7:09 |
| Text token rate | 3-5 tokens per second | 9:12 |
| Audio token rate | more than 10 per second | 9:12 |
| Training audio data volume | about 100 million hours | 20:24 |
| Training data equivalent in human lifetimes | about 200 lifetimes (at 75 years each) | 20:24 |
| Hard drive price change | about three times a year ago | 21:26 |
| Voice model cost comparison | about one-tenth of OpenAI's models | 36:46 |
| Audio model workload relative to language model | about one-half to one-third | 31:40 |
Glossary
- Proactbench
- A benchmark measuring whether a voice model can proactively step into a conversation and judge when it should interrupt.
- Ibench
- A benchmark released by the same team, measuring interruptibility, responsiveness and the match between audio and content.
- back channeling
- Short responses like "mm-hm" and "right" that signal you're listening, especially common in Japanese.
- millennial sandwich
- A communication strategy of praise, criticism, then praise again, used to deliver negative feedback without angering the other person.
- theory of the mind
- A model's ability to infer other people's beliefs, intentions and emotional states.
How to listen
Engineers and founders building voice/multimodal products, especially anyone who needs to judge real-time inference cost, interruption latency and EQ benchmarks.
The first 5 minutes of chit-chat about whether ChatGPT's voice mode is any good can be skipped.