The world is too loud. Read what matters.

No Priors

Diffusion Isn't Just for Images — It Will Win Inference for Language Models

Autoregressive inference must generate one token at a time sequentially, making it a memory-bandwidth bottleneck; diffusion models process multiple tokens in parallel, and the inference workload looks more like training, so Inception is betting on it.

Diffusion ModelsInference EfficiencyModel ArchitectureVoice AgentsAI Startups
A diffusion model pioneer explains why parallelism at inference becomes the architectural battleground, and how a 50-person company finds differentiation in the cracks between giants.

The argument · tap a timestamp to hear it

3:10

Diffusion started because GANs were too unstable

Ermon says that when he was working on image generation, he was very dissatisfied with GANs: they worked, but training was extremely unstable and results were hard to reproduce. So in 2019 he and his PhD students turned to score-based generative models, with the idea of training a neural network to denoise — if it can denoise, it understands image structure, and can then reverse the process to progressively refine clean images from pure noise. This is the underlying technology of modern diffusion models. He stresses that today's best image, video, some music, and many protein models are all based on diffusion.

— Stefano Ermon
5:11

The 2024 paper brought diffusion to text

Ermon's team published a paper in 2024 that, for the first time, got diffusion models to match autoregressive models at GPT-2 scale (under a billion parameters): same perplexity, same degree of data fit, but about 10x faster generation, because diffusion outputs multiple tokens at once. He says this is still academic, but it was enough for him to decide to start a company and scale the technology to commercial scale.

— Stefano Ermon
8:14

Autoregressive inference is a memory-bandwidth bottleneck

Ermon draws an analogy from the 2017 RNN-to-transformer switch to the inference stage: during training, transformers won by processing multiple tokens in parallel, but at inference, autoregressive models are still sequential — before generating the 10th token, all previous tokens must be generated. This workload is extremely memory-bandwidth limited, with most time spent moving weights between memory hierarchies, very little arithmetic, and it's unfriendly to GPUs. Diffusion models at inference look more like training — processing multiple tokens in parallel — so they naturally match the compute patterns GPUs excel at.

— Stefano Ermon
10:16

Inference efficiency also caps RL post-training

Ermon points out that economics are determined by how much intelligence you get per watt and per dollar. Much of the progress in reasoning models comes from test-time compute scaling, and one bottleneck in RL post-training is generating rollouts — letting the model explore, scoring trajectories, and then improving the model. So inference efficiency directly gates post-training efficiency: a model that scales better at inference automatically scales better during post-training. This is the core logic behind his bet on diffusion LLMs, and what he calls ‘more parallel approaches will eventually win’.

— Stefano Ermon
12:19

Mercury targets Haiku, Flash, mini series

Ermon says Inception's Mercury model benchmarks against speed-optimized models like Haiku, Flash, mini nano, with comparable quality but significantly faster, and is already serving customers in production. He stresses that you can't just run a diffusion LLM on vLLM or SGLang; you must build your own serving engine to handle the complexity of real production workloads. This is both a barrier and a moat: even if someone else trains a diffusion LLM, without a corresponding serving engine they can't use it.

— Stefano Ermon
17:21

Voice customer switched from custom chips back to Nvidia GPUs

Ermon gives the example of voice agent company OpenCall: the voice pipeline includes ASR, an inference LLM for tool calls, and TTS, where speed is critical. OpenCall originally served the LLM on Cerebras custom chips to achieve the required speed, but later switched to Inception's diffusion LLM because they could get the same speed on Nvidia GPUs — higher availability, lower cost, better quality. Ermon says once you get used to fast models you can't go back, like broadband.

— Stefano Ermon
25:26

Diffusion is inherently more controllable than autoregressive

Ermon says diffusion models are generally easier to control than autoregressive models: autoregressive must wait until the entire object is generated to know whether it satisfies constraints, is aligned, or what reward the reward function gives; diffusion is coarse-to-fine generation, so you can judge early whether the direction is right and use external reward functions or constraint sets to guide generation. There is considerable evidence in the academic literature supporting this, and some guidance methods are simply impossible with autoregressive models. He sees this as a differentiation space for future product experiences.

— Stefano Ermon
29:29

20% to 30% of tasks are extremely latency-sensitive

Ermon estimates based on OpenRouter's task categories (research, chat, coding, software engineering, log processing, etc.) that about 20% to 30% of tasks are very latency-sensitive. As a lower bound, this portion of the workload can be served by the model that provides the highest quality within a given latency budget. He also admits that we're not yet at frontier intelligence levels, and many workloads still require frontier-level intelligence, so the addressable market for diffusion models is layered.

— Stefano Ermon
35:36

Flash Attention and DPO both came from his lab

Ermon responds to the skepticism that ‘academia can't do large models’: early diffusion work beat GANs at academic scale and was more stable, later growing into Stable Diffusion and Midjourney; he also co-advised Flash Attention, and DPO originated from a project in his group, now widely used to align LLMs and diffusion models. He believes academia's advantage is the ability to make counterintuitive bets, having excellent students, and people not being afraid to take risks — many important AI ideas originated in academia.

— Stefano Ermon

In their own words · checked verbatim

the computation is one left to right one token at a time. You cannot generate the 10th token until you've generated everything that comes before it. That kind of workload is um does not map well to GPUs. that kind of workload is extremely memory bound.

Stefano Ermon8:14

the equivalent at inference time is a diffusion based is a diffusion model because a diffusion model is built to have at inference time a workload where you process many tokens at the same time.

Stefano Ermon9:16

the bitter lesson is that the more parallel solution is the one that is eventually going to win.

Stefano Ermon10:16

once you get used to a fast model, it's hard to go back. It's kind of like broadband, right?

Stefano Ermon17:21

whenever you train these models you're effectively trying to identify common structure by trying to find an efficient way of compressing the data. And so the more you can compress the data, the more structure, the more patterns you're identifying.

Stefano Ermon22:24

right now uh the the the human ingenuity is still like super important and the ability to come up with the the right ideas and kind of like prune the space and and kind of like identify directions that are more promising has been really important to us.

Stefano Ermon33:33

Figures

Diffusion model text generation speed improvementAbout 10x that of an autoregressive model of the same scale5:11
Inception company sizeAbout 50 people, founded about two years ago13:19
Estimated share of tasks extremely sensitive to latency20% to 30%29:29

Glossary

diffusion model
A generative model that starts from pure noise and progressively denoises and iteratively refines a result.
autoregressive model
A model that predicts the next token one by one, generating sequentially from left to right.
perplexity
A metric measuring how well a language model identifies structure in data; lower is better.
serving engine
The software stack that deploys a trained model as an online service and handles production requests.
RL post-training
Continuing to optimize a model after pre-training using reinforcement learning, often requiring many rollouts.

How to listen

Who it's for

Engineers and founders focused on inference cost and latency-sensitive applications (voice agents, coding), plus AI investors who want to understand the architectural competitive landscape.

Skip

The academic reminiscing from 0:08–2:54 can be fast-forwarded; start at 5:11 for the architecture argument.