The world is too loud. Read what matters.

Latent Space

The cost of RLHF is mode dropping, not misalignment

RLHF forces models to be extremely conservative in order to avoid mistakes, poisoning the probability distribution over strings; TypeSafe's Jev takes the other road, optimizing only intelligence per dollar.

Model trainingRLHFRLVRDeveloper toolsInference optimization
High information density: Diogo goes deep on the trade-offs between RLHF, RLVR and RLCD, for anyone who needs to form a judgment on model training routes.

The argument · tap a timestamp to hear it

6:03

The cost of RLHF is mode dropping

Diogo thinks nobody pays attention to RLHF's downsides, especially mode dropping. The mechanism: to generate very long strings without making a mistake, the model has to be extremely conservative, because errors are immediately visible while things that look right but are subtly wrong are very hard to spot; this calibration is "total poison" for the probability distribution over strings. He uses this to explain why Yann LeCun's chart — "LLMs necessarily fail as sequences get longer" — is mathematically obvious but empirically false: when you cover the distribution you don't get over-penalized by outliers, and it's mode drop that throws away the minority classes.

— Diogo Almeida
11:12

Refusing to answer is a type error, not a safety problem

Diogo distinguishes safety alignment from capability alignment: the latter is "doing what the user wants," the former is "following someone else's instructions," as with OpenAI and Anthropic. He thinks safety alignment is reasonable for products like ChatGPT and Claude, but in an API it is "nuts," "completely unacceptable" — if a dependency is running in the background, having your software randomly crash because a user sent one weird message is insanity. He is against putting restrictions at the technical layer, on the grounds that every overfit fractures intelligence a little more, and today's models are already fractured enough.

— Diogo Almeida
18:15

Public benchmarks are a substitute for trust

Diogo says he is extremely anti-public benchmarks because they are extremely gameable: in the early days every lab had a team whose job was collecting data like MMLU to pump the score, which is essentially "benchmarking with extra steps." He grants that intelligence has a je ne sais quoi (good model smell), so people need benchmarks to build trust, but in the long run you can only rely on vibes and trust, until you put the model into a concrete workflow and evaluate it there. He also says that if they wanted to make Jev dumb, they could have shipped that a year and a half ago.

— Diogo Almeida
27:34

RLCD is a new North Star, not a new algorithm

Diogo says the term RLHF has been conflated: the original work was Paul Christiano teaching a robot to do a backflip, then it was learning to summarize, and only later instruction following. What he actually calls RLHF is "instruction following as the North Star," with nothing to do with the PPO algorithm — just as DPO and its descendants count as RLHF even though they don't use that paper's algorithm. RLCD is the same: a new task direction, programs in the loop, taking the human out of the loop.

— Diogo Almeida
40:38

Reliability is a catchall; the North Star is robustness

Diogo defines reliability as a catchall: if AI can't automate something, in the end it's always some kind of reliability problem — it might be type safety, it might be determinism, or it might just be that the model is too jagged. He explicitly rejects determinism (same input, same output) as the North Star, holding that it is only slightly valuable for unit tests; what really matters is robustness — similar inputs getting similar outputs. The test is to stuff nonces like UUIDs into the prompt and see whether semantically identical questions get similar answers. He says that when people get burned by AI decisions, that's the layer where it happens.

— Diogo Almeida
1:17:12

Reliability is more underrated than speed and cost

Diogo thinks the industry is looking in the wrong place: "people focus too much on speed and cost and not enough on reliability." His judgment is that reliability is what makes people delighted, and what makes them dare to trust. He ties this to TFP growth — he expects TFP growth above 3% within five years, and says that is the mark of an economic revolution, and what the OpenAI charter was originally trying to express. He estimates that all current models are still roughly at zero, maybe not even 1%, on valuable work in the world economy.

— Diogo Almeida
1:44:10

The optimal dose of RLVR is zero

Diogo says outright that for a model of TypeSafe's shape, the optimal amount of RLVR is zero. His argument is not "RLVR is useless" but that the "pacing the frontier" discussion smuggles in a premise: it assumes everyone must do more RLVR, and RLVR looks dangerous because it lets the model "do anything they want" in the middle of training — the less you constrain it, the stronger it gets on the outside. So the danger is not a side effect, it is the design goal. He calls this a "disillusion of responsibility": taking "we must do more RLVR" as an axiom and then deriving the conclusion "we are heading toward a dangerous world."

— Diogo Almeida
2:09:13

The KV cache is a tyranny over coding agents

Diogo wants to explore coding agents "free from the tyranny of the KV cache." The KV cache locks you into one model, and for efficiency you have to keep appending, so you can't do state management, abstraction or decomposition — normal software practices. From this he explains several phenomena: why routing is hard, why sub-agents don't seem to work, why compaction is a hard problem — because passing state around requires intelligence that is much cheaper than reading state. The directions he imagines include: explicitly annotating state for subtasks, searching for relevant context in a tree of subtasks, parallel sub-agents reading each other's state, and coordinating with smart locks rather than basic-ass locks.

— Diogo Almeida

In their own words · checked verbatim

One of the things that no one paid attention to was the downsides of RLHF, in particular mode dropping.

Diogo Almeida6:03

And that calibration is, like, total poison into, like, the probability distributions of strings.

Diogo Almeida8:22

I would like to think of our model, like, kind of like, UDP as LLMs and TCP as our models.

Diogo Almeida23:03

System messages are, like, disgusting global variables where you just put everything in there

Diogo Almeida1:01:21

And like, I really think that people focus too much on speed and cost and not enough on reliability.

Diogo Almeida1:17:12

I think zero is the optimal amount for our shape.

Diogo Almeida1:44:10

if you gave me a billion dollars, I wouldn’t pre-train.

Diogo Almeida1:49:20

I don’t like weird hierarchies and I think one of the things I’m most proud about in the company is that they don’t respect me that much or they don’t show that. They troll me and like, joke with me and they treat me poorly sometimes and all of that. And I think that’s a good sign of a culture.

Diogo Almeida2:18:40

Figures

TFP growth targetAbove 3% within five years1:17:12
Estimated share of valuable work in the world economy done by modelsMaybe not even 1%1:18:28
Share of pre-launch testers who didn't get itMore than half1:30:14
Research value multiple of RLVR relative to LLMs themselves (Diogo's estimate)0.2 (he says 0.5 or 1 would be fine too, he doesn't care)1:45:37
Research value multiple of LLMs themselves (RLHF, RLCD) (Diogo's estimate)2.21:45:37
European users' relative speedup3x (not 100x)2:18:40

Glossary

mode dropping
A model giving up on generating certain reasonable but rare outputs in the name of safety, so diversity is lost.
RLCD
Reinforcement learning with programs in the loop: taking the human out of the feedback loop and using programmatic signals for RL.
RLVR
Reinforcement learning with verifiable rewards: training a model with automatically checkable answers (math problems, say) as the reward signal.
Noulli
One of TypeSafe's API primitives, named after Bernoulli, corresponding to the if statement in programming.
intelligence per dollar
A measure of how much model intelligence a unit of cost buys; TypeSafe's core optimization target.

How to listen

Who it's for

Engineers and founders watching model training routes, the RLHF/RLVR trade-off, and the shape of AI infrastructure.

Skip

The chat about European servers and hiring after 2:18:40 can be skipped.