Post-training prowess determines who wins the inference business
AI spending flows primarily to inference rather than training, but what truly determines inference efficiency and cost is post-training optimization—this is why training and inference should be run by the same company.
The video won't play here. Listen to the audio instead:
The argument · tap a timestamp to hear it
Building your own model isn't about protecting against labs, it's about control
Yash explicitly rejects framing ‘owning your own intelligence’ as protection against lab misconduct—he believes labs are staffed by good people and bound by contracts. The real reason is flexibility and control: where models run, how they deploy, whether optimization targets cost, latency, or domain-specific performance. This narrative only emerged once open-source models became genuinely useful. The more immediate risk is practical: models have repeatedly been removed from products, especially in programming; and as labs increasingly enter vertical markets, whether they'll continue serving customers who directly compete with them remains uncertain.
— YashPost-training can beat frontier models, but only when data is genuinely out-of-distribution
Post-training can exceed frontier model performance only if a company's data is genuinely out-of-distribution compared to what labs trained on. But most companies doing post-training aren't after capability gains at all—they're using cheaper models to match frontier performance, optimizing for price-to-performance rather than capability ceiling. The more pressing question: how long can a domain's data remain out-of-distribution? Labs are now saturating programming; other domains will follow.
— YashContinuous learning is bottlenecked by data efficiency, not algorithm design
Pretraining has poor data efficiency but strong generalization; supervised fine-tuning has better data efficiency but limited gains; reinforcement learning with verifiable rewards trades quantity for quality using reproducible environments and validators, yet remains data-inefficient. Yash is blunt: the true blocker for ‘continuous learning’ isn't algorithm choice—it's how to learn efficiently from sparse rewards, a problem that remains unsolved.
— YashNot every company needs pre-training, but all companies have reasoning worth modeling
Responding to Satya's ‘as many models as companies’ framing, Yash doesn't believe every company should pre-train or mid-train its own model, but does believe every company has proprietary decision loops worth capturing. Pharma is the clearest example—each company focuses on different conditions and accumulates distinct experimental data; building your own model is natural. But with banking, Jacob immediately pushed back: how different can one bank's financial data really be from another's? Yash didn't fully engage the objection, instead appealing to different processes and risk preferences, and cited accounting firms—businesses that naturally fragment but then consolidate—as a counterexample.
— Yash / JacobGuard your evaluation sets as fiercely as you guard departing employees
Reinforcement learning works even in unverifiable domains when you convert judgment calls into scoring rubrics—provide expert answers and score against them, and the results are surprisingly good. Different companies find different experts, so their trained models diverge. This makes ‘evaluation sets’ themselves a competitive asset: any public benchmark eventually gets gamed, so Yash argues companies should guard their evaluation sets with the same fierceness they prevent core employees from joining competitors, not share them with labs.
— YashPost-training, not infrastructure, is the real moat in inference
Applied Compute screens customers for just two things: either capability gains are exceptionally valuable, or inference volume is especially large—because the real spend is inference, not training. The more crucial insight: the largest, most expensive workloads should post-train first, because efficiency gains scale as a power law to everything else. Training choices even dictate inference architecture—to enable better parallel tool calls, inference needs prefill-decode separation in its deployment.
— YashThe AI hyperscaler's stack is an inverted pyramid, hardest at the base
Yash positions the company as AWS/GCP/Azure for the GPU era: not just inference—which is currently the first sticky layer in the AI software stack—but the full stack of training, inference, routing, safety monitoring, and agent sandboxing. He compares the structure to an inverted pyramid: at the base, training, which only a handful can do; above it, inference, routing, context and harness engineering, with each layer having lower barriers to entry and more practitioners. Sixteen months old, the company deliberately started with training and is now expanding to inference.
— YashReinforcement learning gains don't generalize across domains like pretraining does
Empirically, capabilities trained on math datasets transfer almost nowhere when moved to law—reinforcement learning gains lack pretraining's cross-domain generalization. Yash uses a metaphor: intelligence is more like a bumpy sphere, where nearby points generalize to each other, but distant regions barely connect at all, nothing like pretraining's holistic reach.
— YashIn their own words · checked verbatim
But I do think the idea of owning your intelligence really is about flexibility and control.
Yash2:06
I basically think like what's blocking continual learning is basically like extremely data efficient training from sparse rewards, which is an unsolved problem.
Yash13:32
Like you wouldn't, your employees are not fungible. You would not like be comfortable with them going to another company and doing work there. Same thing with these models.
Yash26:11
And so, so what we, what our view is, is that the most post training actually wins inference.
Yash29:15
Our big bet was, hey, post training is a massively underlooked part of the stack.
Yash35:26
intelligence is some sort of sphere or circle. It's got jagged points on it. The points that are closer together, you'll see more generalization there.
Yash42:33
It's like water. They find the path of least resistance.
Yash56:09
Figures
| Applied Compute founding to interview | 16 months | 36:29 |
| Online training customer base | tens of thousands to hundreds of thousands of users | 13:32 |
Glossary
- Out-of-distribution (OOD) data
- Data types unseen during model training that determine whether post-training can deliver genuine capability gains.
- RL with verifiable rewards (RLVR)
- Training models on tasks with clear right answers using validators, trading data quantity for quality.
- TCO
- Total cost of ownership; here, the full cost of continuous post-training and online training for a model.
- Harness
- Tool-calling, prompt engineering, and context work layered around a model—does not include modifying the model itself.
- Hyperscaler
- A provider of end-to-end infrastructure like AWS, GCP, or Azure.
How to listen
Founders and enterprise AI leaders deciding whether to build or fine-tune their own models, and investors tracking the business models of inference infrastructure.
Around 4:35, a discussion of the OpenAI-Baseten partnership has low information density and transcription gaps; you can skip it.