Robots don't need to be humanoid to fold laundry or make coffee
Pi uses three main papers to push the robot brain from capability to generalisation to performance; Keli Yiming (柯丽一鸣) argues humanoid form is not the problem worth solving right now, real-robot data is irreplaceable, and the experience data robots generate themselves is the next step.
The video won't play here. Listen to the audio instead:
The argument · tap a timestamp to hear it
Pi's three main papers each answer one question
Keli Yiming breaks Pi's main publications into three keywords: π0 is capability, π0.5 is generalisation, π*0.6 is performance. π0 did three tasks — folding laundry, folding boxes, tabletop pick-and-place — and the November 2024 laundry folding was at a level nobody had seen at the time. π0.5 has to answer the in domain and out of domain question: if the model only performs well in houses it was trained on, its impact is very limited, so they went to a hundred Airbnbs to collect data and found that once the data reaches a certain size, performance in new houses plateaus — you don't need to collect every home in the world. π*0.6 returns to the criticism of "can do everything but nothing well", emphasising simplicity of method and letting the agent collect its own experience data in the real world and put it back into the training pool.
— Ke LiyimingData robots generate themselves is cheaper than human teleoperation
Real-robot data is expensive, and Keli Yiming breaks it into several parts: you need a real robot platform, you need to put it in someone's home, you need to maintain it, you need to find someone to do the tasks, and you need to communicate with that person to do planning — the most complicated part is managing that person. But if you swap the source of data manipulation for an already-trained large model and let the model roll out on its own, the data cost can drop a lot. He is fairly bullish in believing that robots should eventually be cheap and convenient, and that many of today's data worries are actually unimportant over the long arc of time, because all the data collected by so many robots during deployment can be used by me. That is also the meaning of the asterisk in π*0.6: put finished data back into the model.
— Ke LiyimingBad data is good data, but the robot can also generate it itself
Imitation learning has a serious problem called compounding error: every tiny mistake keeps amplifying, and the robot reaches a terrible situation it never imagined and that a human would never touch when collecting data, and once it gets in there it doesn't know how to correct its path. So when collecting data you need not only perfect human data but also correction data on how the robot recovers after entering a bad situation and how it continues to complete the task. But Keli Yiming stresses again that correction data can also come from the robot running on its own — that is the more essential thing about reinforcement learning. On the real-robot versus simulation data debate, he is a black cat white cat as long as it catches the mouse kind of person, but at root he believes in real-robot data, because for tasks like laundry folding with flexible materials, friction and stickiness, a simulator may simply not be able to produce it.
— Ke LiyimingEvaluation is why the robot frontier is hardest to define
In NLP, generating a passage and having people score it was already not easy; robots are worse: you have to run it on the robot to know what it did, to score it, and you are constrained by mechanical physics — some variation you don't even know about will affect model performance. Take the task of picking up a cup: the cup placed anywhere on the table, different lighting, different backgrounds, table height, or even not a table but a pillar, the cup's own angle — endless initial conditions all affect the final performance. Difficulty of evaluation directly means the robotics field, unlike other fields, has no leaderboard where you release a giant dataset, every model runs it, and you see who is number one. So the frontier is very hard to define right now, everyone's direction is scattered, everyone is heading in different directions, and nobody knows which line is the main road.
— Ke LiyimingIf Pi did humanoids, he wouldn't have joined
Keli Yiming talked with Pi within a week of its founding in 2024, asking whether they would do humanoids first and what commercialisation they would do. The atmosphere he felt was that they indeed would not do humanoids. He says if they had said humanoids, he wouldn't have come, because he worried that people from different backgrounds doing humanoids would shift the research focus to how to build a better humanoid robot, whereas his research focus is doing better tasks. He strongly believes you can do very good tasks without a humanoid, and do them faster without a humanoid — folding laundry, folding boxes, making coffee were all done first, and can be done without a humanoid. He thinks this is a matter of taste, and there are endless tasks waiting to be done.
— Ke LiyimingReinforcement learning is essentially Pavlov's dog
Keli Yiming says reinforcement learning is explained by Pavlov's dog: the dog does something, you give it a reward, and from then on it knows this thing is a good thing to do more of; through reward and punishment you make the agent take better actions. It has a lot to do with human self-improvement, and the pursuit of excellence is a very strong factor in reinforcement learning — just practising nonstop. Another important part is exploration, which can be big or small: it might be shifting a muscle slightly to the left for better tennis performance, or it might be switching research direction. And there is attribution: a pet does a series of things and you finally give one reward, and it needs to know which thing was the decisive factor in getting this good reward. π*0.6 contains exploration: collecting so much deployment data, exactly which step in which piece of data was done well and made the whole trajectory get the key reward.
— Ke LiyimingChina's hardware dominance leaves the US unsure how to catch up
Keli Yiming says a year or two ago, wanting to see a demo at the level of the Spring Festival Gala was just a thought, and seeing it actually done was very exciting — behind it are both algorithmic improvements and hardware strength. Looking at the China-US industry comparison at this stage, the most direct feeling is that China has leadership and dominance in hardware; it is hard to imagine a robotics company assembling a robot with not a single Chinese component inside. China's advantage in supply chain and manufacturing is simply too large, and he doesn't even know how Americans would catch up if they tried. At the same time, hardware iteration gives existing algorithms some advantage on the path to commercial deployment. But he also admits that lacking the supply chain and manufacturing link will slow down the iteration speed of US companies.
— Ke LiyimingPremature commercialisation cost Covariant general-purpose generalisation
Keli Yiming explains why Pi is for now not considering commercialisation at all, saying there is a historical reason. Peter Abbeel was Sergey Levine's advisor and has had two startup experiences: one was a system for grading student homework, now used by basically all American universities; the other was founding the robotics company Covariant in 2015 or 2016, to build a general machine learning solution for robots. But at some point in the company's development it began to go deep into logistics and warehouse robots, and looking back from the perspective of large model development, this was actually a distraction from developing large models — because of premature commercialisation it lost general-purpose generalisation, put energy into many commercialisation-related things, and did not really go back to the root questions. So Pi, influenced by this experience, stresses not thinking too much about commercialisation, and everyone very purely saying they want to do research and make its performance the best.
— Ke LiyimingIn their own words · checked verbatim
The first is π0, its keyword should be capability. The second is π0.5, its keyword is generalisation. And the most recent one is π*0.6, its keyword is performance.
第一篇是派灵 它的关键词应该就是能力 第二篇是派0.5 它的关键词是泛化 然后最近的这篇是派0.6星 它的关键词是表现
Ke Liyiming1:55:10
But I think, if the price of the hardware comes down, then in this process, personally I think the current big cost is actually managing the person to collect the data you imagine. But if you can set aside the source of data manipulation and instead put in a large model you have already trained, and let the large model run here, then your data cost should be able to drop a lot.
但我觉得 就是如果机械的价格 降下来的话 就在这个过程中 我个人认为 现在的比较大的大头 其实是管这个人 去收到你想象的这个数据 但如果你可以把 就是数据的操纵源 这个给放在一边 而是换上你 已经训好的一个大模型 让大模型在这里跑 其实你数据的成本 应该是能降很多
Ke Liyiming2:02:14
So the frontier is still very hard to define right now. I think you can see some algorithms, you can see the performance of some algorithms — how should I put it — maybe it's that there is a most core part that everyone can see, and then as it moves toward a more frontier direction it becomes more and more messy, and you need to find a clear line out of the mess.
所以现在这个frontier 还是很难定义的 我觉得可以看到一些算法 可以看到一些算法的表现 就是怎么说呢 可能是 就是一个最核心的部分 大家都看得见 然后随着它向着 更前沿的方向迈进以后 就变得越来越有些乱 就需要你能从分乱中 找到一个清晰的线
Ke Liyiming2:09:17
The atmosphere I felt at the time was that indeed we don't do humanoids. Because if they had said humanoids, I wouldn't have come, you know.
我当时感受到的氛围 就是确实我们不做人型 因为他们要说做人型 我就不来了 你知道吧
Ke Liyiming2:32:30
I think reinforcement learning is a very essential kind of problem, that is, how a person becomes better through experience. This problem, from the most traditional textbook explanation of reinforcement learning, is actually explaining Pavlov's dog.
我觉得强化学习是一个非常本质的一类问题 也就是说一个人如何通过体验变得更好 这个问题从最传统教科书上来说 解释强化学习 其实就是解释这个巴布洛夫的狗
Ke Liyiming2:47:33
It's hard to imagine a robotics company assembling a robot with not a single component inside that is Chinese. I think that's pretty much impossible.
就是很难想象 有一家机器人公司 它 组装好了一台机器人 然这里面没有一个零件 是中国的 我觉得是不太可能
Ke Liyiming3:16:50
Because of premature commercialisation it lost this kind of general-purpose generalisation, and instead put energy into many commercialisation-related things, and didn't really go back to the root questions.
就因为过早的去商业化 而失去了就是这种通用繁华性 反而把精力放在了很多 这种商业化的相关事情上 嗯 没有真的去追溯本源的问题
Ke Liyiming3:17:50
Figures
| Pi valuation | over $5 billion | 0:00 |
| Number of homes for π0.5 data collection | about 100 homes | 1:58:12 |
| Number of Pi employees | about 70 | 2:28:29 |
| Keli Yiming's PhD duration | 7 years | 2:42:32 |
| Record PhD graduation at UW | 9 years | 2:43:33 |
| Productivity gain from Cloud Agent | about 3-4x | 2:31:30 |
| Anhui province rate of admission to Tsinghua/Peking per 10,000 people | about 3-4 people | 25:22 |
| Rate of admission to Tsinghua/Peking per 10,000 people in big cities | about 80 people | 25:22 |
| 2017 ShadowHand price | $500,000 to $1,000,000 | 51:45 |
| 2017 Franka price | tens of thousands of dollars | 52:46 |
Glossary
- VLA / vision-language-action model
- A robot architecture that unifies vision, language and action into a single model.
- in domain / out of domain
- The difference in a model's performance inside and outside the coverage of its training data.
- Diffusion Policy
- A method that uses diffusion models to generate robot actions.
- sim2real
- Transferring a policy trained in a simulator onto a real robot.
- throughput
- The evaluation metric proposed in π*0.6, measuring the amount of successful task completion per unit of time.
- FAST / action representation space
- A Pi paper studying how to learn a better action representation for a large model to predict.
How to listen
Founders, investors and engineers watching embodied intelligence and robot brains, especially those who want to understand Pi's technical route, the real-robot data debate, and the hardware gap between China and the US.
The first 40 minutes on growing up, writing fiction and wuxia games can be fast-forwarded; start from the robot factions and Pi research in the middle.