Diffusion Isn't Just for Images — It Will Win Inference for Language Models
Autoregressive inference must generate one token at a time sequentially, making it a memory-bandwidth bottleneck; diffusion models process multiple tokens in parallel, and the inference workload looks more like training, so Inception is betting on it.
The video won't play here. Listen to the audio instead:
The argument · tap a timestamp to hear it
Diffusion started because GANs were too unstable
Ermon says that when he was working on image generation, he was very dissatisfied with GANs: they worked, but training was extremely unstable and results were hard to reproduce. So in 2019 he and his PhD students turned to score-based generative models, with the idea of training a neural network to denoise — if it can denoise, it understands image structure, and can then reverse the process to progressively refine clean images from pure noise. This is the underlying technology of modern diffusion models. He stresses that today's best image, video, some music, and many protein models are all based on diffusion.
— Stefano ErmonThe 2024 paper brought diffusion to text
Ermon's team published a paper in 2024 that, for the first time, got diffusion models to match autoregressive models at GPT-2 scale (under a billion parameters): same perplexity, same degree of data fit, but about 10x faster generation, because diffusion outputs multiple tokens at once. He says this is still academic, but it was enough for him to decide to start a company and scale the technology to commercial scale.
— Stefano ErmonAutoregressive inference is a memory-bandwidth bottleneck
Ermon draws an analogy from the 2017 RNN-to-transformer switch to the inference stage: during training, transformers won by processing multiple tokens in parallel, but at inference, autoregressive models are still sequential — before generating the 10th token, all previous tokens must be generated. This workload is extremely memory-bandwidth limited, with most time spent moving weights between memory hierarchies, very little arithmetic, and it's unfriendly to GPUs. Diffusion models at inference look more like training — processing multiple tokens in parallel — so they naturally match the compute patterns GPUs excel at.
— Stefano ErmonInference efficiency also caps RL post-training
Ermon points out that economics are determined by how much intelligence you get per watt and per dollar. Much of the progress in reasoning models comes from test-time compute scaling, and one bottleneck in RL post-training is generating rollouts — letting the model explore, scoring trajectories, and then improving the model. So inference efficiency directly gates post-training efficiency: a model that scales better at inference automatically scales better during post-training. This is the core logic behind his bet on diffusion LLMs, and what he calls ‘more parallel approaches will eventually win’.
— Stefano ErmonMercury targets Haiku, Flash, mini series
Ermon says Inception's Mercury model benchmarks against speed-optimized models like Haiku, Flash, mini nano, with comparable quality but significantly faster, and is already serving customers in production. He stresses that you can't just run a diffusion LLM on vLLM or SGLang; you must build your own serving engine to handle the complexity of real production workloads. This is both a barrier and a moat: even if someone else trains a diffusion LLM, without a corresponding serving engine they can't use it.
— Stefano ErmonVoice customer switched from custom chips back to Nvidia GPUs
Ermon gives the example of voice agent company OpenCall: the voice pipeline includes ASR, an inference LLM for tool calls, and TTS, where speed is critical. OpenCall originally served the LLM on Cerebras custom chips to achieve the required speed, but later switched to Inception's diffusion LLM because they could get the same speed on Nvidia GPUs — higher availability, lower cost, better quality. Ermon says once you get used to fast models you can't go back, like broadband.
— Stefano ErmonDiffusion is inherently more controllable than autoregressive
Ermon says diffusion models are generally easier to control than autoregressive models: autoregressive must wait until the entire object is generated to know whether it satisfies constraints, is aligned, or what reward the reward function gives; diffusion is coarse-to-fine generation, so you can judge early whether the direction is right and use external reward functions or constraint sets to guide generation. There is considerable evidence in the academic literature supporting this, and some guidance methods are simply impossible with autoregressive models. He sees this as a differentiation space for future product experiences.
— Stefano Ermon20% to 30% of tasks are extremely latency-sensitive
Ermon estimates based on OpenRouter's task categories (research, chat, coding, software engineering, log processing, etc.) that about 20% to 30% of tasks are very latency-sensitive. As a lower bound, this portion of the workload can be served by the model that provides the highest quality within a given latency budget. He also admits that we're not yet at frontier intelligence levels, and many workloads still require frontier-level intelligence, so the addressable market for diffusion models is layered.
— Stefano ErmonFlash Attention and DPO both came from his lab
Ermon responds to the skepticism that ‘academia can't do large models’: early diffusion work beat GANs at academic scale and was more stable, later growing into Stable Diffusion and Midjourney; he also co-advised Flash Attention, and DPO originated from a project in his group, now widely used to align LLMs and diffusion models. He believes academia's advantage is the ability to make counterintuitive bets, having excellent students, and people not being afraid to take risks — many important AI ideas originated in academia.
— Stefano ErmonIn their own words · checked verbatim
the computation is one left to right one token at a time. You cannot generate the 10th token until you've generated everything that comes before it. That kind of workload is um does not map well to GPUs. that kind of workload is extremely memory bound.
Stefano Ermon8:14
the equivalent at inference time is a diffusion based is a diffusion model because a diffusion model is built to have at inference time a workload where you process many tokens at the same time.
Stefano Ermon9:16
the bitter lesson is that the more parallel solution is the one that is eventually going to win.
Stefano Ermon10:16
once you get used to a fast model, it's hard to go back. It's kind of like broadband, right?
Stefano Ermon17:21
whenever you train these models you're effectively trying to identify common structure by trying to find an efficient way of compressing the data. And so the more you can compress the data, the more structure, the more patterns you're identifying.
Stefano Ermon22:24
right now uh the the the human ingenuity is still like super important and the ability to come up with the the right ideas and kind of like prune the space and and kind of like identify directions that are more promising has been really important to us.
Stefano Ermon33:33
Figures
| Diffusion model text generation speed improvement | About 10x that of an autoregressive model of the same scale | 5:11 |
| Inception company size | About 50 people, founded about two years ago | 13:19 |
| Estimated share of tasks extremely sensitive to latency | 20% to 30% | 29:29 |
Glossary
- diffusion model
- A generative model that starts from pure noise and progressively denoises and iteratively refines a result.
- autoregressive model
- A model that predicts the next token one by one, generating sequentially from left to right.
- perplexity
- A metric measuring how well a language model identifies structure in data; lower is better.
- serving engine
- The software stack that deploys a trained model as an online service and handles production requests.
- RL post-training
- Continuing to optimize a model after pre-training using reinforcement learning, often requiring many rollouts.
How to listen
Engineers and founders focused on inference cost and latency-sensitive applications (voice agents, coding), plus AI investors who want to understand the architectural competitive landscape.
The academic reminiscing from 0:08–2:54 can be fast-forwarded; start at 5:11 for the architecture argument.