The cost of RLHF is mode dropping, not misalignment
RLHF forces models to be extremely conservative in order to avoid mistakes, poisoning the probability distribution over strings; TypeSafe's Jev takes the other road, optimizing only intelligence per dollar.
The video won't play here. Listen to the audio instead:
The argument · tap a timestamp to hear it
The cost of RLHF is mode dropping
Diogo thinks nobody pays attention to RLHF's downsides, especially mode dropping. The mechanism: to generate very long strings without making a mistake, the model has to be extremely conservative, because errors are immediately visible while things that look right but are subtly wrong are very hard to spot; this calibration is "total poison" for the probability distribution over strings. He uses this to explain why Yann LeCun's chart — "LLMs necessarily fail as sequences get longer" — is mathematically obvious but empirically false: when you cover the distribution you don't get over-penalized by outliers, and it's mode drop that throws away the minority classes.
— Diogo AlmeidaRefusing to answer is a type error, not a safety problem
Diogo distinguishes safety alignment from capability alignment: the latter is "doing what the user wants," the former is "following someone else's instructions," as with OpenAI and Anthropic. He thinks safety alignment is reasonable for products like ChatGPT and Claude, but in an API it is "nuts," "completely unacceptable" — if a dependency is running in the background, having your software randomly crash because a user sent one weird message is insanity. He is against putting restrictions at the technical layer, on the grounds that every overfit fractures intelligence a little more, and today's models are already fractured enough.
— Diogo AlmeidaPublic benchmarks are a substitute for trust
Diogo says he is extremely anti-public benchmarks because they are extremely gameable: in the early days every lab had a team whose job was collecting data like MMLU to pump the score, which is essentially "benchmarking with extra steps." He grants that intelligence has a je ne sais quoi (good model smell), so people need benchmarks to build trust, but in the long run you can only rely on vibes and trust, until you put the model into a concrete workflow and evaluate it there. He also says that if they wanted to make Jev dumb, they could have shipped that a year and a half ago.
— Diogo AlmeidaRLCD is a new North Star, not a new algorithm
Diogo says the term RLHF has been conflated: the original work was Paul Christiano teaching a robot to do a backflip, then it was learning to summarize, and only later instruction following. What he actually calls RLHF is "instruction following as the North Star," with nothing to do with the PPO algorithm — just as DPO and its descendants count as RLHF even though they don't use that paper's algorithm. RLCD is the same: a new task direction, programs in the loop, taking the human out of the loop.
— Diogo AlmeidaReliability is a catchall; the North Star is robustness
Diogo defines reliability as a catchall: if AI can't automate something, in the end it's always some kind of reliability problem — it might be type safety, it might be determinism, or it might just be that the model is too jagged. He explicitly rejects determinism (same input, same output) as the North Star, holding that it is only slightly valuable for unit tests; what really matters is robustness — similar inputs getting similar outputs. The test is to stuff nonces like UUIDs into the prompt and see whether semantically identical questions get similar answers. He says that when people get burned by AI decisions, that's the layer where it happens.
— Diogo AlmeidaReliability is more underrated than speed and cost
Diogo thinks the industry is looking in the wrong place: "people focus too much on speed and cost and not enough on reliability." His judgment is that reliability is what makes people delighted, and what makes them dare to trust. He ties this to TFP growth — he expects TFP growth above 3% within five years, and says that is the mark of an economic revolution, and what the OpenAI charter was originally trying to express. He estimates that all current models are still roughly at zero, maybe not even 1%, on valuable work in the world economy.
— Diogo AlmeidaThe optimal dose of RLVR is zero
Diogo says outright that for a model of TypeSafe's shape, the optimal amount of RLVR is zero. His argument is not "RLVR is useless" but that the "pacing the frontier" discussion smuggles in a premise: it assumes everyone must do more RLVR, and RLVR looks dangerous because it lets the model "do anything they want" in the middle of training — the less you constrain it, the stronger it gets on the outside. So the danger is not a side effect, it is the design goal. He calls this a "disillusion of responsibility": taking "we must do more RLVR" as an axiom and then deriving the conclusion "we are heading toward a dangerous world."
— Diogo AlmeidaThe KV cache is a tyranny over coding agents
Diogo wants to explore coding agents "free from the tyranny of the KV cache." The KV cache locks you into one model, and for efficiency you have to keep appending, so you can't do state management, abstraction or decomposition — normal software practices. From this he explains several phenomena: why routing is hard, why sub-agents don't seem to work, why compaction is a hard problem — because passing state around requires intelligence that is much cheaper than reading state. The directions he imagines include: explicitly annotating state for subtasks, searching for relevant context in a tree of subtasks, parallel sub-agents reading each other's state, and coordinating with smart locks rather than basic-ass locks.
— Diogo AlmeidaIn their own words · checked verbatim
One of the things that no one paid attention to was the downsides of RLHF, in particular mode dropping.
Diogo Almeida6:03
And that calibration is, like, total poison into, like, the probability distributions of strings.
Diogo Almeida8:22
I would like to think of our model, like, kind of like, UDP as LLMs and TCP as our models.
Diogo Almeida23:03
System messages are, like, disgusting global variables where you just put everything in there
Diogo Almeida1:01:21
And like, I really think that people focus too much on speed and cost and not enough on reliability.
Diogo Almeida1:17:12
I think zero is the optimal amount for our shape.
Diogo Almeida1:44:10
if you gave me a billion dollars, I wouldn’t pre-train.
Diogo Almeida1:49:20
I don’t like weird hierarchies and I think one of the things I’m most proud about in the company is that they don’t respect me that much or they don’t show that. They troll me and like, joke with me and they treat me poorly sometimes and all of that. And I think that’s a good sign of a culture.
Diogo Almeida2:18:40
Figures
| TFP growth target | Above 3% within five years | 1:17:12 |
| Estimated share of valuable work in the world economy done by models | Maybe not even 1% | 1:18:28 |
| Share of pre-launch testers who didn't get it | More than half | 1:30:14 |
| Research value multiple of RLVR relative to LLMs themselves (Diogo's estimate) | 0.2 (he says 0.5 or 1 would be fine too, he doesn't care) | 1:45:37 |
| Research value multiple of LLMs themselves (RLHF, RLCD) (Diogo's estimate) | 2.2 | 1:45:37 |
| European users' relative speedup | 3x (not 100x) | 2:18:40 |
Glossary
- mode dropping
- A model giving up on generating certain reasonable but rare outputs in the name of safety, so diversity is lost.
- RLCD
- Reinforcement learning with programs in the loop: taking the human out of the feedback loop and using programmatic signals for RL.
- RLVR
- Reinforcement learning with verifiable rewards: training a model with automatically checkable answers (math problems, say) as the reward signal.
- Noulli
- One of TypeSafe's API primitives, named after Bernoulli, corresponding to the if statement in programming.
- intelligence per dollar
- A measure of how much model intelligence a unit of cost buys; TypeSafe's core optimization target.
How to listen
Engineers and founders watching model training routes, the RLHF/RLVR trade-off, and the shape of AI infrastructure.
The chat about European servers and hiring after 2:18:40 can be skipped.