The biggest bottleneck for robots isn't the algorithm, it's the data
Tan Jie has spent ten years doing robotics at Google DeepMind, and his judgment is: the robot foundation model is still not an independent discipline, just a finetune of a multimodal large model; the real bottleneck is data, and collecting data by teleoperation is not the endgame.
The video won't play here. Listen to the audio instead:
The argument · tap a timestamp to hear it
The robot foundation model is not yet an independent discipline
There is a view in China that the embodied intelligence foundation model should not be seen as an extension of large models but as an independent direction. Tan Jie's response is direct: so far not yet. His reasoning is that the vast majority of people doing VLA and robot large models are essentially doing finetuning on multimodal large models — a multimodal model can take in images and language and output text, but it lacks robot action output, so what everyone is doing is completing that output capability. He thinks it could become a more independent discipline only when it hits various bottlenecks, needs a different data format, and needs a more complete world model — but so far there has been no qualitative change.
— Tan JieLanguage models hit a data wall; robots have no data
Tan Jie identifies data as the biggest problem for robots. Language model data is free: there is already a massive amount of language data online, Wikipedia and books can all be digitalized, and language is a narrow domain relative to robotics. But robotics operates in a very complex, unstructured environment where anything can happen, and it needs an extremely large and very diverse amount of data — and that data does not exist today. He is certain that the current data volume is nowhere near enough to saturate the model's capability, so the second bottleneck — say, model architecture — has not even been discovered yet. He describes a data pyramid: at the bottom is internet data, above that egocentric video, above that simulation data, and at the very top robot-specific data; pretraining needs a large amount of low-quality data to learn physical intuition, and only post-training needs a small amount of high-quality data.
— Tan JieCross-embodiment transfer solves the shortage of data
The second breakthrough in Gemini Robotics 1.5 is cross embodiment transfer. Tan Jie says robot data is scarce: data collected on robot A can only be used to learn tasks on A, and if A is upgraded with a different camera position or one more degree of freedom, the previously collected data is useless; data from robots with different configurations cannot be used interchangeably either. They tested three robots — the simple dual-arm Aloha, the more industrial Biarm Franca, and the Optimus humanoid — put the data together, and developed motion transfer technology so that a task seen on robot A can also be executed by robot B. His example: suppose you learned to drive, I never learned, but I also learned to drive. This fundamentally solves the problem of insufficient data, because data collected by any robot can be used by other robots.
— Tan JieThe fast-slow dual model is a transition, not the endgame
Gemini Robotics 1.5 splits the system into an ER (embody reasoning, slow thinking) and a VLA (fast thinking, producing robot action). Tan Jie sees this as a transitional approach, not the ultimate one. The reason is that it is currently constrained by compute and model size: slow thinking has to do reasoning and web search, requiring a very large model, which makes it hard to make decisions five to ten times per second; only a small model can make decisions five to ten times per second. To unify into one model, it would have to be very large because it has to do reasoning, and current compute is not enough to support that. He also points out that the two models currently communicate in language, and language is not a high-bandwidth communication method — it loses a lot of information — so the ideal is a single model with no extra interface inside and no information loss.
— Tan JieThe definition of simulation data is being rewritten by generative models
Tan Jie says the definition of simulation is becoming increasingly blurry. Simulation used to mean physics simulation — Bullet, MojoCo, Isaac Gym — solving physics equations in a computer to compute motion trajectories. Now, with the rise of video generation models, many people think simulation is just generating a video, and as long as it looks physically correct, that is also simulation in a new sense. He judges that in the not-too-distant future, traditional physics-simulation will gradually be replaced by generative-model simulation. On economics, he actually says generating video may be more expensive because it requires compute cost, but it solves the scene generation problem: traditional simulation requires hand-building 500 home scenes one by one, while a generative model just needs 500 different prompts. He also points out that the current bottleneck is hallucination and non-physical phenomena — for example, when a person is generated doing gymnastics, you don't know how many legs will come out.
— Tan JieOnly what can be played is a world model; static video is not
Tan Jie gives his definition of a world model: if you are given the previous frame and then the robot's action, you can predict the next frame. By this standard, the vast majority of current video generation models are not world models, because they cannot change the next frame by inputting a robot action or a human action. He gives a comparison: VIO is a video generation model, but Genie is more like a world model, because Genie can be played — you can change the generated next frame by pressing keys, you can drive a dragon to turn left and right, and every frame has an input that changes the next frame. He says Google DeepMind is the best in this direction, Sora is also good, OpenAI is also good, and many small companies are working on it too.
— Tan JieTactile sensing only became necessary once dexterous hands appeared
Tan Jie recounts his own journey of judgment on touch. He had always believed tactile sensing was important, because humans feel touch through their skin every day. But the Aloha paper from Stanford showed him for the first time that a person could teleoperate a robot with pure vision to do very complex things, including taking a very thin credit card out of a leather bag — he thought such a thing could only be done with touch, and Aloha slapped him hard in the face, so he held that belief for a while, thinking vision could do 95% of things. It was only when dexterous hands became widespread that he went to use teleoperation to control a dexterous hand using scissors — you put your hand through the two loops and you can use the scissors, but without tactile feedback he didn't know when to open and when to close, because the loops are large, and opening and closing at the wrong time just means the hand moves inside the loops without actually controlling the scissors. Only then did he realize that with dexterous hands touch is very important, and that his earlier view that it wasn't important was limited by the hardware of the time.
— Tan JieGoogle Robotics went from ten loose people to an army
When Tan Jie joined nine years ago, the team had only ten people, Google Brain had only a few dozen, management was very loose, and every researcher who came in was a one-man show who did whatever they wanted, and resource allocation was not a problem. He describes himself then as like a very well paid PhD or assistant professor, with autonomy, but personal impact was very limited, and it was hard to gather a group of people to do something big. Now Robotics has 150 people, the March version of the first generation of Gemini Robotics may already have 120 people on the author list, and this time it may be 160 to 180. He says Google realized it did not want to be just an academic lab but very costly, so it created an environment through promotion, performance review, incentives and structure to let more people solve bigger things together. He himself changed along with the team, so he did not have culture shock.
— Tan JieIn their own words · checked verbatim
Whether it is really a very independent, separate discipline — so far not yet
它到底是不是一个非常独立的单独的学科 so far not yet
Tan Jie13:06
Just to give an example, that is, suppose you learned to drive, I have never learned this task of driving, but I also learned to drive — it can transfer across embodiments
就举个例子 就是说假设你学会了开车 我从来没有学过开车这个任务 但我也学会了开车 它可以跨本体
Tan Jie30:21
They believe that in the robotics industry there will soon be a qualitative change, a huge change; if that change happens, I'm on a wrong ship, I must be at the right ship
他们相信在机器人行业 很快会有一个质的变革 巨大的变革 如果这个变革发生 I'm on a wrong ship 我一定要at the right ship
Tan Jie1:36:24
But I think in the era of large models, you really need to burn a lot of money before you can see results
但是我觉得这个在大模型时代 你真的需要烧很多钱 你才能看到结果
Tan Jie1:46:33
I think the vast, vast majority of people, especially those not in the robotics industry, overestimate the development of robots
我觉得绝大数的绝大多数的人 尤其是不是在机器人行业里的人 对机器人的发展是overestimate的
Tan Jie2:00:46
Figures
| Google Robotics team size (nine years ago) | 10 people | 1:24:14 |
| Google Robotics team size (now) | 150 people | 1:24:14 |
| Gemini Robotics first generation March version author list size | about 120 people | 1:26:14 |
| Gemini Robotics 1.5 author list size | 160 to 180 people | 1:26:14 |
| Tan Jie's weekly working hours | 70 to 80 hours | 1:29:15 |
| Success rate on fine manipulation task (zipping a zipper) | 30% to 40% | 16:06 |
| Success rate on simple pick and place | 90-something percent to 100% | 16:06 |
| Expected time for robots to reach GPT-3/GPT-4 level | two to three years | 17:08 |
| Expected time for robots to truly land | an additional five to ten years | 17:08 |
| Unseen-device task completions brought by MSR researcher | 10 of 25 tasks completed | 1:54:37 |
Glossary
- VLA / vision-language-action model
- A model that takes in images and language and directly outputs robot actions; it is the mainstream form of the robot brain today.
- cross embodiment transfer
- Transferring data collected and skills learned on one robot to another robot with a different configuration.
- motion transfer
- The technology in Gemini Robotics 1.5 that lets data from different embodiments be used interchangeably; Tan Jie calls it the secret sauce.
- sim to real
- Training a policy in simulation and then deploying it to a real robot; Tan Jie's first Google paper used this method.
- egocentric video
- Manipulation video shot from a person's point of view, collectable by wearing glasses or a camera; large in volume but with a gap relative to robot morphology.
- embodiment gap
- The degree of morphological difference between two robots; the larger the gap, the harder cross-embodiment transfer is.
How to listen
Founders, engineers and investors working on robotics or embodied intelligence; anyone who wants to know Google DeepMind's real internal judgment on data, architecture and world models.
The first 20 minutes of guest bio and the graphics-to-robotics backstory; fast-forward to the data section at 23:44.