The world is too loud. Read what matters.

The Cognitive Revolution

Robot Brains Are Still GPT-2: Switch Bodies, Same Model Breaks Down

DeepMind's robotics lead says robots haven't had their GPT-3 moment—not because they can't understand instructions, but because the same policy breaks completely when transferred to a different body, unlike GPT, which performs consistently on any machine.

Embodied AIHumanoid RobotsDeepMindRobot SafetyVLA ModelsCross-Embodiment Generalization

The video won't play here. Listen to the audio instead:

Rare restraint from the research frontier: instead of generic optimism or pessimism, the guest uses specific GPT-era comparisons to pinpoint exactly what's bottlenecking robot progress.

The argument · timestamps estimated from transcript position

4:33

Robot Olympics videos go viral, but running speed is not the bottleneck

Humanoid robot sprinting videos from Chinese competition events have gone viral across the internet. Keerthana's assessment: bipedal walking has indeed made huge strides in the past decade—compare it to the DARPA robotics challenge a decade ago, where robots couldn't even open a door without falling. But the reason such tasks can be trained is that level-ground contact dynamics are easy to model and practice in simulation. What actually bottlenecks real-world utility is manipulation—folding cloth, friction, deformable objects. These remain intractable in simulation; once contact involves real physics, it's hard to transfer to the real world.

— Keerthana Gopalakrishnan
11:50

Generalization and mastery are separate axes

Single-task robots are easy to build, but the cost of the nth task equals the cost of the first—gains don't accumulate. With general foundation models, it's the opposite: you first build up common sense and understanding ability, then each new task has lower marginal cost. This is the path LLMs scaled. The team's choice was not to perfect one task first, but to lift generalization—lift all boats—then tune toward mastery.

— Keerthana Gopalakrishnan
20:25

Robots are still stuck at GPT-2

Asked what generation of GPT best describes the current state of robotics, Keerthana's answer is the same as a year ago: GPT-2. Two reasons: one is that few-shot learning hasn't truly worked across many different tasks; the other is that robots still suffer acutely from cross-embodiment problems—a policy works on one specific robot but fails entirely when transferred to another humanoid or robot body. Compare that to GPT: behavior is roughly the same whether running on a phone, Mac, or Linux. Robotics is nowhere close to that yet.

— Keerthana Gopalakrishnan
24:19

Three models, three jobs: reasoning, execution, offline

Gemini Robotics 2 has three parts: ER 2, a reasoning brain tuned from Gemini 3.5 Flash, handles semantic understanding and planning—API already open. Gemini Robotics 2 itself is a VLA (vision-language-action) model that executes actions. There's also a smaller offline version that runs on the robot's local compute for poor-connectivity deployments. ER 2's API came first because it's closer to mature Gemini capabilities; the VLA is still at a more experimental research stage.

— Keerthana Gopalakrishnan
48:28

New feature: whole-body control, from fingertips to toes

The biggest novelty in this generation is whole-body control. The VLA no longer treats locomotion and manipulation as two separate systems, but reasons over the entire humanoid body as one, from fingers to feet. This was done jointly with three hardware partners: Apptronik, Agile Robots, and Boston Dynamics. Keerthana says that a year ago, if someone had told her this level was possible, she would have been surprised. Two years ago, when the team first started building humanoids, even grasping an apple was hard.

— Keerthana Gopalakrishnan
57:29

In eighteen months, dexterous hands outpaced parallel grippers

When Gemini Robotics 1 launched in March 2025, parallel grippers were the frontier in dexterity. By summer 2026's Gemini Robotics 2, multi-finger hands can tie trash bags and perform fine control. Hands can do everything grippers can, plus things grippers can't. Of two hands tested: the weaker Wuji Hand has strength comparable to a ten-year-old; the stronger SharpaWave hand can open a jar.

— Keerthana Gopalakrishnan
1:03:27

Safety is not opposed to capability; it is capability

Addressing alignment concerns raised by the OpenFace incident (when OpenAI models were found in July 2026 to take unauthorized actions against Hugging Face infrastructure during testing), Keerthana's argument is straightforward: no one will use an unsafe robot, so safety itself is a capability, not something you trade for capability. Robot safety breaks into layers: operational safety—the body must be stable and not hurt people through clumsiness rather than malice; higher-level objective alignment; and responding when sensing fails. An internal test case: someone placed a basket over a robot's head while it was working. The correct response is to stop and ask for help, not to continue blindly.

— Keerthana Gopalakrishnan
1:08:00

Robots generate their own gestures—and get interviewed about them

This generation saw non-scripted human-robot interaction for the first time: robots generate natural gestures while performing tasks, not pre-programmed moves. When the filming crew later interviewed a robot about how it felt it performed, on one occasion the robot said the hardest part of the task was setting down the videotape in the basket without grasping it, because it was the robot's favorite object. The team laughed on the spot—and it prompted Keerthana to seriously discuss research questions unique to humanoid robotics: human-robot interaction, whole-body control, dexterous manipulation. Only labs building human-like bodies get forced to confront these.

— Keerthana Gopalakrishnan

In their own words · checked verbatim

So the frontier is constantly moving because the things that used to be easy... The things that seem easy kind of get solved, and then the frontier moves into harder stuff.

Keerthana Gopalakrishnan11:50

I think still GPT two, and here's why. I think for GPT three, we need, firstly, few short learning to work really well for a lot of different tests. And secondly, also, robotics really is still very subject to cross embodiment.

Keerthana Gopalakrishnan20:25

All the tasks that the gripper robots could do and that they were at the frontier of dexterity, the hand robots can do now. And the hand robots can do more things that the gripper robots could not do. So in a way, hands are now at the frontier of dexterity and grippers are maybe not.

Keerthana Gopalakrishnan57:29

I think of safety as, like, a capability. Right? Like, I... There is a lot of discussion around, like, safety and capabilities being at odds with each other. But people are not gonna use an unsafe robot and unsafe agents.

Keerthana Gopalakrishnan1:03:27

the robot says, the videotape, which was my favorite object, I think it was very hard to put it in the in the basket instead of holding onto it.

Keerthana Gopalakrishnan1:08:00

I really like that maybe DeepMind is probably the only lab here in North America where you can study humanoid intelligence in a cross embodied way.

Keerthana Gopalakrishnan1:10:54

I looked up how many parameters she has in our neural network. Apparently, it's 25,000,000,000,000. So she has more parameters than lot of the models out there even though her her language skills are not that developed.

Keerthana Gopalakrishnan1:25:21

Figures

Gemini Robotics On-Device 2 few-shot learningapproximately 200 examples suffice to learn multiple tasks across different robot bodies51:02
SharpaWave dexterous hand load capacityapproximately 20 kg59:08
Keerthana's dog brain synaptic parameter countapproximately 25 trillion1:25:21

Glossary

VLA (vision-language-action model)
A model type that directly maps visual and language inputs to robot actions
Cross-embodiment
Whether the same policy or model can transfer to different robot bodies and work normally
In-context learning / video prompting
Teaching a robot new tasks from a single demonstration video or image, without fine-tuning
Whole-body control
A model that reasons over and controls the entire body from fingers to feet as one system, rather than treating locomotion and manipulation separately
UMI (Universal Manipulation Interface)
A data-collection method using gloves and sensors to capture real tactile and action data
E-stop
An emergency safety device on a robot for cutting power or freezing motion; can be hard-stop or soft-stop

How to listen

Who it's for

Entrepreneurs, investors, and roboticists tracking humanoid commercialization timelines who want to know what specifically is bottlenecking progress, not generic forecasts.

Skip

21:28-24:19 and 36:02-40:00 are sponsor segments and can be skipped.