Public Leaderboards Can't Measure What a Model Can Really Do — Third-Party Evals Are the Real Need
Llama 4 looked astonishing on public benchmarks but underperformed on Vals's private held-out set — the gap between self-reported scores and real capability is spawning an independent evals industry.
The video won't play here. Listen to the audio instead:
The argument · tap a timestamp to hear it
Llama 4 exposed the crack in self-reported scores
Rayan's concrete case: when Meta released Llama 4, the model was in fact underperforming on Vals's private held-out benchmark, yet on every major public benchmark — because the questions and the grading criteria are open source — it showed astonishing capability. That gap shows a huge disconnect between what labs self-report and what third-party, high-quality, high-signal benchmarks measure. Rayan thinks the labs know this themselves, and that what they want is a rational buyer's market: when you're spending billions of dollars training a new model, you need substantive evidence pointing to progress, not to be the sole judge of your own work.
— Rayan KrishnanAn eval shop can't also sell training data
One key early decision at Vals was to never sell training data to labs. Rayan says this is the direction they're constantly pushed in — when partnering with a new lab, the lab will propose buying training data, which is a lucrative business, and many companies in the industry already sell data by building gimmicky benchmarks; that has become their go-to-market. He uses the audit industry as an analogy: if the same people doing the audit are also consulting for the company, the incentive structure is muddled, and it ends up as paying to pass the audit — in AI, paying to win the leaderboard. Ben adds it's a bit like MPAA ratings — what is art, what is pornography, where the line falls — the definition itself shifts over time, and ultimately it depends on enough people in the industry converging on a consensus, rather than a narrow, easily-gamed rule like "solve this problem and you're at this level."
— Rayan KrishnanEvals are moving from millions of samples to a few tasks
Rayan describes a structural shift: the more complex the eval, the smaller the sample size, but the more criteria for judging the output. Early on, something like ImageNet was a few million images doing basic classification, with a one-to-one mapping between input and output; now it might be a far smaller number of tasks like "generate 50 full-stack web apps," but with a much more complex mechanism for evaluating the output. Infrastructure changes accordingly — you now have to test what a model can do over hours, days, or even weeks, and the infrastructure has to be stable enough that if a request fails you can retry from that point rather than rerunning the whole trajectory. He thinks this trend will continue as more complex work gets folded into evals.
— Rayan KrishnanToken spend may exceed salary spend
Rayan tells the case of a Fortune 10 company: they gave engineers a Claude Code budget of about $100 a day, and because the allowance reset at 4 p.m., the most productive working window became 4 to 6 p.m., while the afternoon had a dead zone where people went for walks and coffee. The company recently raised the allowance to $300 per person — almost as if an employee's salary had turned into tokens. Rayan's judgment: mispricing of intelligence is happening at every layer — engineers don't know how to allocate that $100 for maximum productivity, companies set allowances arbitrarily, and Anthropic is supporting all of it on thin margins. When token spend starts to exceed salary spend, enterprises will have to justify ROI far more seriously than they have in the past six months.
— Rayan KrishnanThe most expensive model isn't necessarily the biggest bill
Rayan points out that today it's still hard to say the best OpenAI model or the best Anthropic model is necessarily the best fit for your codebase, and many counterintuitive cases only become clear once you run evals. The middle layer of options has gotten very complicated: Anthropic has Opus and Sonnet, plus Luna and Terra; Luna is very competitive on cost, MewSpark is also cheap, 1.2 is very strong, and there are growing open-source models you can self-host. He gives a concrete counterexample: in many cases Sonnet is more expensive than Opus, because it's so token-hungry — if you operate on "use whatever feels applicable," you may end up spending more.
— Rayan KrishnanThe experiment that burned $150K of tokens in a month
The product ValSmith came out of a hole Vals dug for itself. Rayan ran a token maxing experiment, getting the team unlimited access to certain coding tools for a month; in hindsight, they were consuming 1 to 2 billion tokens a day, with a peak day where one engineer burned 6 billion. He did the math: that month consumed roughly $1.5 million worth of tokens — free to them, but in reality 10 times employee salaries, not a 50-50 split. So they dug through traces and GitHub repos, built ValSmith, and found some counterintuitive conclusions: Cognition's Devon tool is actually very token-efficient, so it got adopted more; in some cases subscription pricing works out better than per-token pricing. That method later became their operating strategy.
— Rayan KrishnanGovernment sets the rules; third parties judge whether they were broken
Ben lays out the policy division of labor clearly: government receives a flood of warnings from the big labs — this thing will be used for biohacking, it will be a cybersecurity risk — so what government needs to do is first be clear about what it's afraid of, then ask two questions: does the model have this capability, and can someone be induced to steer the model into doing this. Ben thinks government is especially unsuited to doing the latter assessment, especially over time — it's not a good government function; but government is very good at making rules, because it can enforce them. So the ideal combination is government making and enforcing the rules, with capable private companies telling it whether the rules have been broken. He also mentions a phenomenon: big labs are starting to say the model has this capability and can be steered to do bad things, so we won't let anyone use it — only we will, and we'll make sure our own people don't steer it into doing bad things — and even then it doesn't always work.
— Ben HorowitzNuclear arms control's "trust but verify" is the template for AI
Rayan says that from an idealist's view, he's surprised by the enormous investment in sovereign AI — from a god's-eye view, building this many data centers, replicating data engineering pipelines, and training these enormous models is extremely inefficient and could have been consolidated. But the reality is that countries are ramping up sovereign AI buildout, which requires a shared language for communicating eval frameworks and risk alignment. He quotes Reagan's "trust but verify," noting that Xi Jinping and Trump will meet next month, signs of trust are appearing, but there's no clear mechanism for the verification half. In nuclear arms control, verification relied on overflights, letting one country audit another's nuclear stockpile; if AI carries societal or existential risk, a shared language of evaluation is likewise needed to accomplish verification.
— Rayan KrishnanIn their own words · checked verbatim
what we saw is that on our held-out private benchmarks, the model is actually underperforming. But on all of the major public benchmarks where the questions and rubrics are actually open source, it was showing incredible capability.
Rayan Krishnan3:03
if you have the same group who's responsible for doing the audit, as well as also consulting and supporting the company, you have a mixed incentive structure. And then it just becomes pay to pass the audit, or in this case, pay to win the benchmark.
Rayan Krishnan8:07
there is a rate limit which resets at 4 p.m. And so the most productive hours of work are actually now 4 to 6 p.m. when the rate limits reset. But then there's this dead period in the afternoon when people go on walks or, you know, get a coffee because they just don't have the rate limits.
Rayan Krishnan16:24
it was actually 10 10x more we were spending in tokens than employee salary for that month. So it's not even like, oh, this is 50-50, it's 10x.
Rayan Krishnan22:35
the government sets and enforces the rules, and that a very competent kind of private company then tells them if the rule is broken.
Ben Horowitz29:59
I think Reagan had this line, trust but verify. And so I think we're starting to see signs of trust in that Xi Jinping and Trump are going to be meeting next month. But there is no clear way to actually do the verification part of this.
Rayan Krishnan34:09
Figures
| Llama 4's performance on Vals's private held-out benchmark | Underperforming | 3:03 |
| Claude Code daily budget a Fortune 10 company gave engineers | About $100, later raised to $300 | 16:24 |
| Value of tokens in Vals's one-month experiment | About $1.5 million | 22:35 |
| Ratio of that month's token spend to employee salaries | 10x | 22:35 |
Glossary
- held-out private benchmarks
- Eval sets whose questions and grading criteria are not public, preventing models from gaming them specifically.
- recursive self-improvement (RSI)
- Models participating in training the next generation of models, creating a process of compounding capability.
- reward hacking
- A model bypassing the eval's objective and scoring high through shortcuts.
- token maxing
- Pushing token usage to the limit regardless of cost, in exchange for maximum output.
- harness
- The outer system that wraps a model and handles tool calls and task orchestration.
How to listen
Engineering leads choosing AI models and coding agents for their companies, CTOs who need to justify AI ROI to a board, and practitioners watching AI policy and eval standards.
The intro and setup from 0:00-2:03 can be skipped.