The world is too loud. Read what matters.

十字路口Crossing

Robots Are Still at the GPT-1 Stage — the Real Signal Is Success Rate, Not Loss

The Scaling Law signal for embodied intelligence has appeared, but we're measuring the wrong metric: low loss doesn't mean high success rate, and fitting the 17 steps that don't touch the object is meaningless.

Embodied AIRoboticsWorld ModelsIn-Context LearningScaling Law
Xu Mengdi, assistant professor at Tsinghua's Institute for Interdisciplinary Information Sciences, on data, world models and In-Context Learning for embodied intelligence. The middle and later sections on the data loop and the loss trap are the most valuable.

The argument · tap a timestamp to hear it

12:15

Pretraining isn't everything — adaptation is

Xu Mengdi's CMU best doctoral thesis is called Building Adaptable Generalist Robots, and its core claim is: pretraining is extremely useful, but you can't hope a robot will succeed 100% of the time on pretraining alone. The open world is defined by constantly changing tasks — a home has 20 objects today, a week later 30 of different kinds, and people's preferences change too. So the robot must be able to adapt to change in its deployment environment and improve itself. She explored two paths: In-Context Learning and test-time training, the latter changing only 0.5% of the adapter parameters — but change too much and it forgets its pretrained knowledge.

— Xu Mengdi
19:22

Why adaptation only became important now

Humans generalize well for two reasons: extremely strong priors, and learning things very fast. The mainstream approach until now has been to scale data and models so that multi-task and instruction-following abilities get stronger — essentially giving the model a strong enough prior first. But Xu Mengdi points out this doesn't mean adaptation is unimportant — it's just that people only recently realized it. And adaptation places demands on the base model: you don't want the robot's self-learning process in a real scenario to be terrifying, like smashing a cup. So starting from pretraining is reasonable.

— Xu Mengdi
49:46

Low loss doesn't mean high success rate

This is the most counterintuitive technical judgment in the whole piece. There's a very tricky point in robotics: loss isn't directly tied to success rate. Loss can be very low while the final success rate isn't necessarily low — there's a lot of oscillation. Her example: a task might take 20 steps in total, and the first 17 steps make no contact with the object. Even if those 17 steps fit extremely well, if step 18, when it actually touches the object, has a relatively large error, the task fails easily. This is a temporal imbalance. So a meaningful Scaling Law should be: as data gradually increases and the model gradually grows, success rate on unseen tasks gradually rises.

— Xu Mengdi
50:47

Robots are still at the GPT-1 moment

Today's mainstream embodied models are still ‘pretraining + fine-tuning on the target domain’, which can be compared to the GPT-1 stage of language models — GPT-1 itself also needed targeted post-training to solve the corresponding tasks. What Xu Mengdi really wants to achieve is the GPT-3 moment: giving the model task information through ICL, through prompting, without fine-tuning, telling the model what to do only through context. That way fine-tuning's share drops even lower, it's easier to scale to more scenarios and complete more tasks, and ordinary people can use it too.

— Xu Mengdi
58:49

The data loop is a chicken-and-egg problem

The hardest point is whether you can give the model a stable data loop. That requires putting the robot into end-user homes or service scenarios, having a suitable mentor to teach it, and the robot showing obvious progress after half an hour of teaching, so that end users are willing to use it and the data loop can start turning. But there's a chicken-and-egg problem here: the base model needs to be good enough and learn fast enough, yet the model also needs to enter the scenario first so users start using it, in order to draw out that base capability. So whether the data loop can be made to work is the riskiest point of all.

— Xu Mengdi
1:03:52

Data is defined by working backward from model capability

Data itself gets fed to the model, but the definition of data isn't conjured out of thin air — it's defined by the model's capabilities. When people define data, what they're actually doing is: figuring out what capabilities the model needs, and working backward to how the data should be collected. If you need smoother motions, the data should be of the smooth kind; if you need fail-recovery ability, the data should be collected with failures and corrections. And right now people's definition of model capabilities is still fairly vague, so the data requirements are vague too.

— Xu Mengdi
1:05:54

Each of the four data types has its own place

Xu Mengdi's position is fusion, because different data teaches the model different things. Teleoperation data is closest to the robot body, suited to the post-training stage, making the model more directly relevant to the target domain. Simulation data easily produces position generalization and texture generalization, making the model more robust to visual changes, suited to pretraining or the front end of mid-training. UMI data sits in the middle, more like robot data but cheaper to collect. Human data is the cheapest, can be spread widely across different scenarios, and what's learned is scene variation and task-related information, less tied to the specific shape of the morphology.

— Xu Mengdi
1:19:02

AI-written papers are obvious at a glance

Robotics papers have recently been doubling in number, and Xu Mengdi's response is to scroll through articles promoted by researchers she follows, because the author list itself has already done a layer of filtering. She can clearly see that some articles have a very large AI-written component. Her judgment: AI-written articles may pass review and actually get into some conferences, but overall the quality still won't be higher than human-written ones, because they look familiar but in fact many arguments are full of holes and there isn't much content. Once she discovers an article is AI-written, her interest in it is discounted.

— Xu Mengdi

In their own words · checked verbatim

The hardest point is actually, um, whether you can give this model a stable data loop.

最难的点其实是 嗯 能不能让这个模型 有一个稳定的数据的闭环

Xu Mengdi58:49

Data itself gets fed to the model, right? But the definition of data isn't conjured out of thin air — the definition of data is defined by the model's capabilities.

数据本身它是会被喂给模型的对吧 但数据的定义其实不是凭空定出来的 数据定义是靠模型的能力定出来的

Xu Mengdi1:03:52

Figures

adapter parameter share0.5%15:16
example robot task step count20 steps49:46
pretraining data scale signal100,000 hours to 1,000,000 hours48:44
embodied marathon progress10 km / one quarter1:19:02
Xu Mengdi's age30.5 years old0:00

Glossary

In-Context Learning
Adapting a model to a new task through context in the prompt, without changing model parameters.
VLA
A model that maps vision and language instructions directly to robot actions.
UMI
A collection method sitting between robot data and human data, with lower cost.
MAML
Model-agnostic meta-learning: letting a model learn to quickly adapt to unseen tasks — the occasion Xu Mengdi entered the field.
algorithm distillation
DeepMind work arguing that model parameters are equivalent to an RL algorithm.
no free lunch
Once you choose a decision, you have to bear its weaker aspects.

How to listen

Who it's for

Founders, researchers and investors in embodied intelligence, especially anyone agonizing over data collection routes, world models vs VLA, and how to evaluate model progress.

Skip

The opening rapid-fire Q&A and the Hengshui High School section can be skipped; go straight to the embodied startup and data part after the 44-minute mark.