The world is too loud. Read what matters.

硅谷101

The benchmark maker's red line: don't sell your own test data

Many benchmark creators are tempted to sell data—but if the data contaminates their own benchmarks, they destroy not just their own leaderboard but every model relying on it.

Data industryAI startupsBenchmarkingReinforcement learningRubricsPost-training
Two frontline practitioners break down the complete chain from data annotation through RL environments, evaluation rubrics, and the benchmark business—high information density, plenty of industry insights, worth listening to in full.

The argument · tap a timestamp to hear it

1:01

Billion-dollar data company valuations hide what's actually driving their value

AfterQuery, founded by three twenty-somethings, had a $300 million April A-round that jumped to $3.2 billion by September—setting a YC record for fastest startup-to-unicorn growth. Scale AI and Mercor are already valued over $29 billion and $10 billion respectively. Yet what kinds of data these companies actually sell and why model makers pay such premiums for it remains opaque.

10:08

Expert-written rules let weak models grade stronger ones accurately

A Rubric's core logic is weak-to-strong supervision: an expert codifies their judgment criteria into a scoring framework, and then even a much weaker model, by following that framework, can produce scores matching the expert's. This solves the problem in domains like history or law where ground truth exists but is unstructured—enabling strong reinforcement learning where simple pattern-matching fails.

— Yun Zhong
12:08

Objective benchmarks in any field could replicate coding's AI breakthroughs

ALE's founding insight: the coding domain has numerous public, easy-to-verify benchmarks that drove ‘vibe coding's’ explosive growth. ALE aims to replicate this model across other domains where objective scoring is possible, enabling agents to solve long-horizon tasks with real economic value and verifiable outcomes. Currently it offers over 150 public tasks across 50+ industries; the next version targets over 1,000 tasks.

— Yi You
22:14

Selling benchmark test data destroys the credibility of all dependent models

Both practitioners emphasize that benchmarking and data sales are distinct enterprises that must never mix. The industry has seen data companies cross this line—taking their own benchmark test data and selling it to model makers for training. This doesn't just destroy the benchmark's authority; it undermines the credibility of every other model relying on that leaderboard. It's a line you simply cannot cross.

— Yun Zhong
34:19

Data company value lies in sourcing rare data, not processing it

The data business splits into acquisition and processing. The real difficulty—and where true value lives—is acquisition: buying private domain data, commercial software rights, even game source code from verticals. Whoever solves sourcing for a particular domain gains enormous leverage. This is why even small two- or three-person data companies command high valuations; the bottleneck isn't compute or labor, it's access.

— Yun Zhong
38:20

Life sciences benchmarks cover barely a tenth of the actual landscape

ALE's team mapped biology's workflow down to a third level using Nature's taxonomy, identifying roughly 3,000 fine-grained nodes, then cross-checked how many existing life-science benchmarks cover them. The result: coverage is around 10%, possibly under 5%. There isn't even one comprehensive benchmark for the field, revealing that the data gap in professional domains far exceeds what most expect.

— Yi You
48:24

Even trusted code benchmarks fail due to data contamination and test bugs

This February, OpenAI stopped reporting SWE-bench Verified scores for two reasons: test scripts have defects that can fail correct solutions, and models can replicate the original human-written fixes under certain prompts, indicating pretraining likely exposed them to the relevant content. The benchmark's scores have only climbed from 74% to 80% over the past six months; everything beyond that likely concentrates on these contamination and testing issues. ALE consequently didn't adopt it.

50:24

Realistic fake data stayed hidden until models vastly exceeded the task

ALE has faced real instances of expert data fabrication. Attracted by the chance to earn paper authorships, some contributors simply used agents to synthesize garbage data—the falsification was obvious on inspection. But the harder problem is polished fake data: claims of running a chemical simulation that never actually ran. These may only surface when models' abilities far outpace the task itself, causing bulk answers to coherently contradict ground truth. Reverification by experts is expensive.

— Yi You

In their own words · checked verbatim

AfterQuery, founded by three people in their twenties, had a $300 million valuation when announcing its A-round in April; by September's new round, its valuation had jumped to $3.2 billion.

由三名20多岁的年轻人创办的 AfterQuery 今年4月宣布A轮融资的时候 估值还是3亿美元 到了9月 新一轮融资 就已经把它的估值推到了32亿美元

For weak models to supervise strong ones, the condition is that an expert has written down their knowledge; then a weak model can follow that and grade the answers.

就我能不能用弱模型来监督强模型 那它这里成立的条件就是说 我有个专家把他脑海里的知识 给写下来了 这样弱的模型拿着它改卷子就行了

Yun Zhong10:08

How we scale vibe coding's explosive growth across every other domain.

我们怎么样把 Vibe coding(氛围编程) 迅速发展的这趋势 推广到vibe everything

Yi You12:08

This is a line you absolutely cannot cross—your business cannot contaminate your benchmark's authority. Otherwise you destroy not just your benchmark but everyone else's models too.

当然这个是 大家千万不能碰的一条线 不能让你做的生意 污染你这个benchmark的权威性 不然的话 你不仅毁了你的benchmark 也毁了其他人的模型

Yun Zhong22:14

Even in the code domain, whoever solves the problem of sourcing private code repositories gains an enormous advantage.

就甚至代码领域 其实采购一些私域的代码仓库 谁能把这个问题解决好 其实就是一个巨大的优势

Yun Zhong34:19

They can demand 2–3 million RMB per source; I don't think any data company or benchmark maker has the budget for that.

它要价能要到两三百万人民币一条 我觉得哪家数据公司 或者别人做benchmark 都没有这个财力去做这个事情

Yi You36:20

Under certain prompts, models can reproduce the original human-written fixes for questions, meaning they likely encountered this content during training.

在特定提示下面 模型能够复现部分题目原本的 人类修复代码 也就是说 模型很可能在训练过程当中 就已经接触过相关内容了

He even spun up an agent to synthesize a bunch of garbage data—some of it was obviously fabricated the moment you looked at it.

他甚至起了agent 去合成一堆乱七八糟的数据 甚至有些数据交上来 我就肉眼看出来他就是编的

Yi You50:24

Figures

AfterQuery A-round valuation$300 million1:01
AfterQuery September round valuation$3.2 billion1:01
Scale AI valuation (after Meta investment)over $29 billion1:01
Mercor latest round valuation$10 billion2:02
ALE current public tasksover 150, covering 50+ industries4:06
ALE V2 target tasksover 1,0004:06
Niche game source code price2–3 million RMB36:20
SWE-bench Verified score change (past 6 months)74% to 80%48:24
Life sciences benchmark coverage in ALE V2approximately 10%, possibly under 5%38:20
Mercor manufacturing engineer hourly rate$100–$200 per hour56:27

Glossary

Rubric
An expert-written scoring standard that allows weaker models to grade stronger models' outputs by following fixed rules.
RL environment
A configured setup with tasks, tools, and feedback mechanisms allowing AI to take actions and learn from results.
Weak-to-strong supervision
When experts codify their judgment standards into rules, weaker models can use those rules to accurately evaluate stronger models.
Reward hacking
A model achieving high scores without genuinely solving the task by finding exploitable loopholes in the scoring mechanism.
Bench-maxxing
Optimizing for a specific benchmark score without the improvement necessarily translating to real-world capability gains.
Recursive self-improvement (RSI)
Long-horizon training where AI iteratively improves its own capabilities by researching and refining itself.

How to listen

Who it's for

Founders and investors interested in the AI data annotation, evaluation, and post-training supply chain; engineers concerned with model training data provenance and validation mechanisms.

Skip

Skip the opening about offline events and hackathons; it doesn't affect the main content.