The world is too loud. Read what matters.

张小珺·商业访谈录

The biggest bottleneck for robots isn't the algorithm, it's the data

Tan Jie has spent ten years doing robotics at Google DeepMind, and his judgment is: the robot foundation model is still not an independent discipline, just a finetune of a multimodal large model; the real bottleneck is data, and collecting data by teleoperation is not the endgame.

RoboticsEmbodied AIWorld ModelsDataGoogle DeepMind
High information density; the middle and later sections on the data pyramid, motion transfer, the definition of a world model, and Silicon Valley 996 are the most valuable. The first 20 minutes of guest bio can be fast-forwarded.

The argument · tap a timestamp to hear it

13:06

The robot foundation model is not yet an independent discipline

There is a view in China that the embodied intelligence foundation model should not be seen as an extension of large models but as an independent direction. Tan Jie's response is direct: so far not yet. His reasoning is that the vast majority of people doing VLA and robot large models are essentially doing finetuning on multimodal large models — a multimodal model can take in images and language and output text, but it lacks robot action output, so what everyone is doing is completing that output capability. He thinks it could become a more independent discipline only when it hits various bottlenecks, needs a different data format, and needs a more complete world model — but so far there has been no qualitative change.

— Tan Jie
23:44

Language models hit a data wall; robots have no data

Tan Jie identifies data as the biggest problem for robots. Language model data is free: there is already a massive amount of language data online, Wikipedia and books can all be digitalized, and language is a narrow domain relative to robotics. But robotics operates in a very complex, unstructured environment where anything can happen, and it needs an extremely large and very diverse amount of data — and that data does not exist today. He is certain that the current data volume is nowhere near enough to saturate the model's capability, so the second bottleneck — say, model architecture — has not even been discovered yet. He describes a data pyramid: at the bottom is internet data, above that egocentric video, above that simulation data, and at the very top robot-specific data; pretraining needs a large amount of low-quality data to learn physical intuition, and only post-training needs a small amount of high-quality data.

— Tan Jie
29:21

Cross-embodiment transfer solves the shortage of data

The second breakthrough in Gemini Robotics 1.5 is cross embodiment transfer. Tan Jie says robot data is scarce: data collected on robot A can only be used to learn tasks on A, and if A is upgraded with a different camera position or one more degree of freedom, the previously collected data is useless; data from robots with different configurations cannot be used interchangeably either. They tested three robots — the simple dual-arm Aloha, the more industrial Biarm Franca, and the Optimus humanoid — put the data together, and developed motion transfer technology so that a task seen on robot A can also be executed by robot B. His example: suppose you learned to drive, I never learned, but I also learned to drive. This fundamentally solves the problem of insufficient data, because data collected by any robot can be used by other robots.

— Tan Jie
34:22

The fast-slow dual model is a transition, not the endgame

Gemini Robotics 1.5 splits the system into an ER (embody reasoning, slow thinking) and a VLA (fast thinking, producing robot action). Tan Jie sees this as a transitional approach, not the ultimate one. The reason is that it is currently constrained by compute and model size: slow thinking has to do reasoning and web search, requiring a very large model, which makes it hard to make decisions five to ten times per second; only a small model can make decisions five to ten times per second. To unify into one model, it would have to be very large because it has to do reasoning, and current compute is not enough to support that. He also points out that the two models currently communicate in language, and language is not a high-bandwidth communication method — it loses a lot of information — so the ideal is a single model with no extra interface inside and no information loss.

— Tan Jie
52:41

The definition of simulation data is being rewritten by generative models

Tan Jie says the definition of simulation is becoming increasingly blurry. Simulation used to mean physics simulation — Bullet, MojoCo, Isaac Gym — solving physics equations in a computer to compute motion trajectories. Now, with the rise of video generation models, many people think simulation is just generating a video, and as long as it looks physically correct, that is also simulation in a new sense. He judges that in the not-too-distant future, traditional physics-simulation will gradually be replaced by generative-model simulation. On economics, he actually says generating video may be more expensive because it requires compute cost, but it solves the scene generation problem: traditional simulation requires hand-building 500 home scenes one by one, while a generative model just needs 500 different prompts. He also points out that the current bottleneck is hallucination and non-physical phenomena — for example, when a person is generated doing gymnastics, you don't know how many legs will come out.

— Tan Jie
56:45

Only what can be played is a world model; static video is not

Tan Jie gives his definition of a world model: if you are given the previous frame and then the robot's action, you can predict the next frame. By this standard, the vast majority of current video generation models are not world models, because they cannot change the next frame by inputting a robot action or a human action. He gives a comparison: VIO is a video generation model, but Genie is more like a world model, because Genie can be played — you can change the generated next frame by pressing keys, you can drive a dragon to turn left and right, and every frame has an input that changes the next frame. He says Google DeepMind is the best in this direction, Sora is also good, OpenAI is also good, and many small companies are working on it too.

— Tan Jie
1:09:57

Tactile sensing only became necessary once dexterous hands appeared

Tan Jie recounts his own journey of judgment on touch. He had always believed tactile sensing was important, because humans feel touch through their skin every day. But the Aloha paper from Stanford showed him for the first time that a person could teleoperate a robot with pure vision to do very complex things, including taking a very thin credit card out of a leather bag — he thought such a thing could only be done with touch, and Aloha slapped him hard in the face, so he held that belief for a while, thinking vision could do 95% of things. It was only when dexterous hands became widespread that he went to use teleoperation to control a dexterous hand using scissors — you put your hand through the two loops and you can use the scissors, but without tactile feedback he didn't know when to open and when to close, because the loops are large, and opening and closing at the wrong time just means the hand moves inside the loops without actually controlling the scissors. Only then did he realize that with dexterous hands touch is very important, and that his earlier view that it wasn't important was limited by the hardware of the time.

— Tan Jie
1:24:14

Google Robotics went from ten loose people to an army

When Tan Jie joined nine years ago, the team had only ten people, Google Brain had only a few dozen, management was very loose, and every researcher who came in was a one-man show who did whatever they wanted, and resource allocation was not a problem. He describes himself then as like a very well paid PhD or assistant professor, with autonomy, but personal impact was very limited, and it was hard to gather a group of people to do something big. Now Robotics has 150 people, the March version of the first generation of Gemini Robotics may already have 120 people on the author list, and this time it may be 160 to 180. He says Google realized it did not want to be just an academic lab but very costly, so it created an environment through promotion, performance review, incentives and structure to let more people solve bigger things together. He himself changed along with the team, so he did not have culture shock.

— Tan Jie

In their own words · checked verbatim

Whether it is really a very independent, separate discipline — so far not yet

它到底是不是一个非常独立的单独的学科 so far not yet

Tan Jie13:06

Just to give an example, that is, suppose you learned to drive, I have never learned this task of driving, but I also learned to drive — it can transfer across embodiments

就举个例子 就是说假设你学会了开车 我从来没有学过开车这个任务 但我也学会了开车 它可以跨本体

Tan Jie30:21

They believe that in the robotics industry there will soon be a qualitative change, a huge change; if that change happens, I'm on a wrong ship, I must be at the right ship

他们相信在机器人行业 很快会有一个质的变革 巨大的变革 如果这个变革发生 I'm on a wrong ship 我一定要at the right ship

Tan Jie1:36:24

But I think in the era of large models, you really need to burn a lot of money before you can see results

但是我觉得这个在大模型时代 你真的需要烧很多钱 你才能看到结果

Tan Jie1:46:33

I think the vast, vast majority of people, especially those not in the robotics industry, overestimate the development of robots

我觉得绝大数的绝大多数的人 尤其是不是在机器人行业里的人 对机器人的发展是overestimate的

Tan Jie2:00:46

Figures

Google Robotics team size (nine years ago)10 people1:24:14
Google Robotics team size (now)150 people1:24:14
Gemini Robotics first generation March version author list sizeabout 120 people1:26:14
Gemini Robotics 1.5 author list size160 to 180 people1:26:14
Tan Jie's weekly working hours70 to 80 hours1:29:15
Success rate on fine manipulation task (zipping a zipper)30% to 40%16:06
Success rate on simple pick and place90-something percent to 100%16:06
Expected time for robots to reach GPT-3/GPT-4 leveltwo to three years17:08
Expected time for robots to truly landan additional five to ten years17:08
Unseen-device task completions brought by MSR researcher10 of 25 tasks completed1:54:37

Glossary

VLA / vision-language-action model
A model that takes in images and language and directly outputs robot actions; it is the mainstream form of the robot brain today.
cross embodiment transfer
Transferring data collected and skills learned on one robot to another robot with a different configuration.
motion transfer
The technology in Gemini Robotics 1.5 that lets data from different embodiments be used interchangeably; Tan Jie calls it the secret sauce.
sim to real
Training a policy in simulation and then deploying it to a real robot; Tan Jie's first Google paper used this method.
egocentric video
Manipulation video shot from a person's point of view, collectable by wearing glasses or a camera; large in volume but with a gap relative to robot morphology.
embodiment gap
The degree of morphological difference between two robots; the larger the gap, the harder cross-embodiment transfer is.

How to listen

Who it's for

Founders, engineers and investors working on robotics or embodied intelligence; anyone who wants to know Google DeepMind's real internal judgment on data, architecture and world models.

Skip

The first 20 minutes of guest bio and the graphics-to-robotics backstory; fast-forward to the data section at 23:44.