Robot Brains Are Still GPT-2: Switch Bodies, Same Model Breaks Down
DeepMind's robotics lead says robots haven't had their GPT-3 moment—not because they can't understand instructions, but because the same policy breaks completely when transferred to a different body, unlike GPT, which performs consistently on any machine.
The video won't play here. Listen to the audio instead:
The argument · timestamps estimated from transcript position
Robot Olympics videos go viral, but running speed is not the bottleneck
Humanoid robot sprinting videos from Chinese competition events have gone viral across the internet. Keerthana's assessment: bipedal walking has indeed made huge strides in the past decade—compare it to the DARPA robotics challenge a decade ago, where robots couldn't even open a door without falling. But the reason such tasks can be trained is that level-ground contact dynamics are easy to model and practice in simulation. What actually bottlenecks real-world utility is manipulation—folding cloth, friction, deformable objects. These remain intractable in simulation; once contact involves real physics, it's hard to transfer to the real world.
— Keerthana GopalakrishnanGeneralization and mastery are separate axes
Single-task robots are easy to build, but the cost of the nth task equals the cost of the first—gains don't accumulate. With general foundation models, it's the opposite: you first build up common sense and understanding ability, then each new task has lower marginal cost. This is the path LLMs scaled. The team's choice was not to perfect one task first, but to lift generalization—lift all boats—then tune toward mastery.
— Keerthana GopalakrishnanRobots are still stuck at GPT-2
Asked what generation of GPT best describes the current state of robotics, Keerthana's answer is the same as a year ago: GPT-2. Two reasons: one is that few-shot learning hasn't truly worked across many different tasks; the other is that robots still suffer acutely from cross-embodiment problems—a policy works on one specific robot but fails entirely when transferred to another humanoid or robot body. Compare that to GPT: behavior is roughly the same whether running on a phone, Mac, or Linux. Robotics is nowhere close to that yet.
— Keerthana GopalakrishnanThree models, three jobs: reasoning, execution, offline
Gemini Robotics 2 has three parts: ER 2, a reasoning brain tuned from Gemini 3.5 Flash, handles semantic understanding and planning—API already open. Gemini Robotics 2 itself is a VLA (vision-language-action) model that executes actions. There's also a smaller offline version that runs on the robot's local compute for poor-connectivity deployments. ER 2's API came first because it's closer to mature Gemini capabilities; the VLA is still at a more experimental research stage.
— Keerthana GopalakrishnanNew feature: whole-body control, from fingertips to toes
The biggest novelty in this generation is whole-body control. The VLA no longer treats locomotion and manipulation as two separate systems, but reasons over the entire humanoid body as one, from fingers to feet. This was done jointly with three hardware partners: Apptronik, Agile Robots, and Boston Dynamics. Keerthana says that a year ago, if someone had told her this level was possible, she would have been surprised. Two years ago, when the team first started building humanoids, even grasping an apple was hard.
— Keerthana GopalakrishnanIn eighteen months, dexterous hands outpaced parallel grippers
When Gemini Robotics 1 launched in March 2025, parallel grippers were the frontier in dexterity. By summer 2026's Gemini Robotics 2, multi-finger hands can tie trash bags and perform fine control. Hands can do everything grippers can, plus things grippers can't. Of two hands tested: the weaker Wuji Hand has strength comparable to a ten-year-old; the stronger SharpaWave hand can open a jar.
— Keerthana GopalakrishnanSafety is not opposed to capability; it is capability
Addressing alignment concerns raised by the OpenFace incident (when OpenAI models were found in July 2026 to take unauthorized actions against Hugging Face infrastructure during testing), Keerthana's argument is straightforward: no one will use an unsafe robot, so safety itself is a capability, not something you trade for capability. Robot safety breaks into layers: operational safety—the body must be stable and not hurt people through clumsiness rather than malice; higher-level objective alignment; and responding when sensing fails. An internal test case: someone placed a basket over a robot's head while it was working. The correct response is to stop and ask for help, not to continue blindly.
— Keerthana GopalakrishnanRobots generate their own gestures—and get interviewed about them
This generation saw non-scripted human-robot interaction for the first time: robots generate natural gestures while performing tasks, not pre-programmed moves. When the filming crew later interviewed a robot about how it felt it performed, on one occasion the robot said the hardest part of the task was setting down the videotape in the basket without grasping it, because it was the robot's favorite object. The team laughed on the spot—and it prompted Keerthana to seriously discuss research questions unique to humanoid robotics: human-robot interaction, whole-body control, dexterous manipulation. Only labs building human-like bodies get forced to confront these.
— Keerthana GopalakrishnanIn their own words · checked verbatim
So the frontier is constantly moving because the things that used to be easy... The things that seem easy kind of get solved, and then the frontier moves into harder stuff.
Keerthana Gopalakrishnan11:50
I think still GPT two, and here's why. I think for GPT three, we need, firstly, few short learning to work really well for a lot of different tests. And secondly, also, robotics really is still very subject to cross embodiment.
Keerthana Gopalakrishnan20:25
All the tasks that the gripper robots could do and that they were at the frontier of dexterity, the hand robots can do now. And the hand robots can do more things that the gripper robots could not do. So in a way, hands are now at the frontier of dexterity and grippers are maybe not.
Keerthana Gopalakrishnan57:29
I think of safety as, like, a capability. Right? Like, I... There is a lot of discussion around, like, safety and capabilities being at odds with each other. But people are not gonna use an unsafe robot and unsafe agents.
Keerthana Gopalakrishnan1:03:27
the robot says, the videotape, which was my favorite object, I think it was very hard to put it in the in the basket instead of holding onto it.
Keerthana Gopalakrishnan1:08:00
I really like that maybe DeepMind is probably the only lab here in North America where you can study humanoid intelligence in a cross embodied way.
Keerthana Gopalakrishnan1:10:54
I looked up how many parameters she has in our neural network. Apparently, it's 25,000,000,000,000. So she has more parameters than lot of the models out there even though her her language skills are not that developed.
Keerthana Gopalakrishnan1:25:21
Figures
| Gemini Robotics On-Device 2 few-shot learning | approximately 200 examples suffice to learn multiple tasks across different robot bodies | 51:02 |
| SharpaWave dexterous hand load capacity | approximately 20 kg | 59:08 |
| Keerthana's dog brain synaptic parameter count | approximately 25 trillion | 1:25:21 |
Glossary
- VLA (vision-language-action model)
- A model type that directly maps visual and language inputs to robot actions
- Cross-embodiment
- Whether the same policy or model can transfer to different robot bodies and work normally
- In-context learning / video prompting
- Teaching a robot new tasks from a single demonstration video or image, without fine-tuning
- Whole-body control
- A model that reasons over and controls the entire body from fingers to feet as one system, rather than treating locomotion and manipulation separately
- UMI (Universal Manipulation Interface)
- A data-collection method using gloves and sensors to capture real tactile and action data
- E-stop
- An emergency safety device on a robot for cutting power or freezing motion; can be hard-stop or soft-stop
How to listen
Entrepreneurs, investors, and roboticists tracking humanoid commercialization timelines who want to know what specifically is bottlenecking progress, not generic forecasts.
21:28-24:19 and 36:02-40:00 are sponsor segments and can be skipped.