Faceless running isn't inelegant—it's the optimum reinforcement learning selected
To run faster, Tiangong (天工) Omni automatically abandoned the human intuition of arm-swing balance, instead using waist rotation—the solution the machine learned may not look elegant, but it obeys the physics-constrained optimum.
The video won't play here. Listen to the audio instead:
The argument · tap a timestamp to hear it
Punishment functions forced reinforcement learning to innovate a new gait
Faceless running was not designed—it emerged from the training setup. The team deliberately restricted the robot's shoulder-joint rotation angles to block the intuitive balancing strategy of arm swing. To run fast while maintaining balance under these constraints, the robot trained itself into a different running style using waist rotation instead. Xiong Youjun calls this a rare moment of emergence in embodied AI: the optimization process discovered a solution the team had not conceived.
— Xiong YoujunNobody knows how much data embodied-AI scaling laws will actually require
Tiangong trained its universal motion controller on several hundred hours of human motion capture; NVIDIA's Sonic controller used roughly 700+ hours—comparable scale. But when asked how much data embodied AI's scaling law will need, Xiong freely admits the field doesn't know yet. The gap between current data (hundreds of hours) and plausible requirements (millions of hours) is still orders of magnitude. More provocatively, he doubts whether stacking data alone is even the right path—large language models work this way because the internet supplies infinite text, but embodied AI may need a different learning paradigm, closer to how human children explore and fail.
— Xiong YoujunThe industry jumped to full autonomy in a single year, and computing power drove it
Last year's marathon race fielded 200+ robots; only Tiangong ran fully autonomously; the rest relied on human teleoperation for guidance. This year, most competitors achieved autonomy in all but the occasional obstacle course. Xiong attributes this shift not to accumulated time but to concentrated investment: the entire industry shifted resources and talent toward robot brains—perception, decision-making, learned control. The payoff came fast.
— Xiong YoujunFewer than five independent startups have the compute to train embodied-AI brains
The episode explores an industry claim: outside of government labs and Big Tech, fewer than five venture-backed companies have enough compute to train large brain models. Xiong did not dispute this. He acknowledged that training an embodied AI brain demands extreme computational resources. Tiangong draws on Baidu's cloud infrastructure and their own data-collection centers, which now hold nearly 200,000 hours of high-quality egocentric footage—a rare advantage.
— Xiong YoujunHonda's 2000s robotics paradigm became obsolete when deep learning arrived
Xiong's passion for robotics traces to 2002: a video of Honda's P2 (the ancestor of ASIMO) escorting a Chinese leader through a factory. The vision was electric—human-robot collaboration would define the future. But Honda's entire approach rested on mathematical models, not learning. The robot carried a heavy external compute backpack and had weak strategic reasoning and vision. When deep learning and large language models emerged, that entire technical architecture became obsolete overnight—a Nokia moment for Honda, which lost its pole position in humanoid robotics.
— Xiong YoujunHigh-level teleoperation delegates human judgment to machines without replacing thought
Some competitors lost their entries for lying about autonomy while relying on teleoperation. Yet the rules explicitly permit teleoperation because it operates at different levels. Framewise puppet control is one thing; task-level teleoperation is another. A human can tell the robot a complex goal—eg, "check if that breaker should be flipped"—and the robot decides its own path, navigation, obstacles, and execution. This is collaboration between human judgment and machine capability, not evasion.
— Xiong YoujunNational-team KPIs optimize for industry breakthroughs, not commercial returns
Xiong's move from CTO at Ubtech to head of the Robotics Innovation Center was a shift from company to country: from maximizing commercial gain to finding shared friction across the whole industry. The center has released the Tiangong robots, Huisi Awakening software, and other foundational work as open-source, so that commercial companies don't duplicate low-level infrastructure and waste social resources. The early shareholders—Ubtech, Xiaomi, Yizhuang State-owned Capital—were not chasing returns; they were betting on the technical direction itself.
— Xiong YoujunEvery sensor reading introduces irreversible error that caps robot precision
Large language models work with purely digital I/O: text in, tokens out, negligible loss. Robots are different. Each input—vision, touch, force—and each motor command introduces noise and drift. Unlike software, these errors compound and cannot be fully corrected. Precision degrades with every physical interaction. Xiong confirmed this bottleneck and framed the solution: unify different modalities—vision, language, touch—into a single learned representation space, then reason and act from that shared substrate. This unified model, called Pelican Unified, is what his team is now building.
— Xiong YoujunIn their own words · checked verbatim
We deliberately restricted shoulder-joint rotation angles. To run that fast while maintaining balance, the robot trained itself into a new gait. It's genuine self-learning and self-evolution.
故意是把他肩关节的转动角度 做了一些限制 然后机器他为了跑那么快 他自己需要保持平衡 然后自己训练出跑制 所以他确确实实是自我学习 自我进化的一个结果
Xiong Youjun2:01
We've always thought that relying purely on data collection might not be the right approach.
我们一直觉得完全靠数据去逮集的话 这可能不一定是一个正确的一个方式
Xiong Youjun10:04
Our data collection volume today is already quite large—the first-person footage alone is approaching 200,000 hours—and it's high-quality data.
今天我们的数据采集量其实已经非常大了 连采集的采集量快接近20万小时了 而且这是高质量的数据
Xiong Youjun18:11
Different modalities' signal representations introduce losses, so we're now pursuing unified representation across all modalities.
不同的模态的信号 信息的这个表征 会带来一些损耗 所以其实我们现在也在追求 对多模态的这个信息进行统一的表征
Xiong Youjun1:09:47
Robots that don't meet safety standards cannot be sold. Otherwise they might cause unnecessary harm to society and users. It's a fundamental bottom line.
你机器人达不到安全的标准是不允许销售的 否则的话可能会对社会对人的用户造成一些不必要的一些意外 这是一个底线的问题
Xiong Youjun1:20:56
Figures
| Tiangong Omni 100-meter final race time | 8.64 seconds | 5:01 |
| Tiangong Omni 100-meter rehearsal time | 9.44 seconds | 5:01 |
| Tiangong universal motion controller training data | several hundred hours of human motion | 8:04 |
| NVIDIA Sonic controller training data | 700+ hours of human motion | 8:04 |
| Innovation Center open-source data collected | nearly 200,000 hours | 18:11 |
| Fastest JD Logistics human sorting speed | 2,000+ pieces/hour | 57:34 |
| Chinese robot logistics sorting speed | 1,200–1,300+ pieces/hour | 57:34 |
Glossary
- VLA
- Vision-Language-Action model—end-to-end robot control that unifies visual perception, language understanding, and motor output in a single neural network
- Egocentric capture
- First-person video footage from cameras worn by human operators, used to train robots through imitation learning at low cost
- World model
- A learned model that lets robots internally predict and reason about physical cause-and-effect in their environment
- Unified model
- A single representation space that fuses VLA, multimodal LLM, and world-model capabilities into one coherent reasoning system
- Scaling law
- The principle that model performance improves predictably with greater data and compute—proven for LLMs, still unverified for embodied AI
How to listen
Investors and founders focused on embodied AI and robotics, plus anyone trying to understand the open-source dynamics between 'national team' labs and commercial companies.
0:00–2:01 is intro and small talk; skip straight to the technical breakdown of faceless running.