The world is too loud. Read what matters.

张小珺·商业访谈录

Zhang Ya-Qin: Information Intelligence Tops Out in Five Years, Humanoid Robots Need Another Ten

He splits AGI into three layers — information, physical, biological: information intelligence reaches human level within five years, but that is the IQ of a brain; reading ten thousand books still won't teach you to swim, and physical intelligence has only just begun.

AGIEmbodied intelligenceAutonomous drivingBrain-computer interfaceRobotics
A timetable and classification framework from an academician. Medium information density; the second half, on lifespan and the extension of the species, is more interesting than the first.

The argument · tap a timestamp to hear it

4:04

A first-rate research institute cannot be measured by quantitative review

Zhang Ya-Qin recalls that when Microsoft Research China was setting its goals, the consultant and Nobel laureate Rodger Ely gave a counterintuitive answer: how do we measure whether we are first-rate? "When you are not worrying about this question, when you are not talking about this question, you have become first-rate." He later ran into the same line on ChatGPT — when does the Turing test count as passed? "When you pass it, you'll know." From this he derived three goals for the institute: recruit first-rate talent, set long-term directions, and cultivate first-rate research talent — and for the first two or three years, 80% of the time went into recruiting.

— Zhang Ya-Qin
6:05

Running a research institute today, the hard part isn't hiring but choosing topics

Zhang Ya-Qin says that compared with 1998, doing AIR now is actually "relatively easy" — China already has world-class research results and talent, and within the Tsinghua environment recruiting good teachers and good students is not a problem. The real difficulty has shifted: in the past Microsoft Research did whatever the people it found wanted to do, "they did whatever they liked," and without a leading figure in a field you couldn't get it off the ground; now you can find people in any direction, so you have to choose the direction yourself. AIR therefore chose three vertical directions: intelligent transportation, IoT, and AI plus life sciences — in today's language, robotics, embodied intelligence, biological intelligence, edge intelligence and agents.

— Zhang Ya-Qin
7:06

Foundation models are a company's job, not a school's

Why doesn't AIR do foundational general-purpose large models? Zhang Ya-Qin's reason is the resource structure: foundation models need enormous compute, data and scenarios, they need money poured in, and they need concrete scenarios — hard for a school to do. He compares them to an ecosystem's operating system, spread out horizontally. More crucial is the nature of the innovation — he argues that innovation in foundation models "is not zero-to-one innovation, it is more one-to-one-hundred, even one-hundred-to-N innovation": there is algorithmic innovation in it, but much of it is engineering innovation, which is what companies should be doing, and in the US it is also companies doing it. AIR's approach is to cooperate with enterprises and use their resources to do it.

— Zhang Ya-Qin
13:08

Information intelligence is IQ; read as many books as you like, you still can't swim

Zhang Ya-Qin uses a person as an analogy: information intelligence is a person with an especially clever head, who has read ten thousand books, understands everything, can solve problems, invent formulas, write articles, paint, and knows more than the average person in every field — that is the IQ of a brain, and he judges it can be reached within five years. But "no matter how many books you read, you still can't swim — you still have to go swim," and that is physical intelligence. He goes further: no matter how much a drinker reads or how clever he is, if he doesn't drink it still won't do — things in the physical world must be interacted with in that world. The first thing to land in physical intelligence is autonomous driving, because it is a close problem; humanoid robots are an open problem and need another ten years.

— Zhang Ya-Qin
15:14

Robots come in three kinds, and home robots are the hardest

Zhang Ya-Qin himself divides robots into three scenarios: home robots, industrial robots, social robots. Home is the easiest to understand — caring for the elderly, doing housework — but the tasks are too complex and fragmented, and they involve human safety and household habits; he considers this the hardest category. Social robots are police, security guards, food delivery, driving — running on the streets and interacting a lot with everyone. Industrial robots are the easiest to define, because the goals are determined and the scenarios fixed. He hopes the technology behind all three kinds of robots will be common: 70% to 80% of the back end the same, with different front ends attached on top, and the back end being one huge multimodal large model.

— Zhang Ya-Qin
20:15

Humanoid form is so it can use humanity's infrastructure

Why is a home robot best humanoid? Zhang Ya-Qin gives two reasons. First, interaction: a humanoid robot made to look like a person — you might not even be able to tell — communicates differently, you can have heart-to-hearts with it, it becomes a companion, a butler, and it can help you care for the elderly. Second, infrastructure: our current society, staircases and buttons are all made for people, and a humanoid robot can directly use the existing infrastructure — climb stairs, press buttons. He adds that this is more a matter of choice, and the same goes for social robots — a police officer too small or too big makes you uncomfortable either way, and a humanoid form can avoid the uncanny valley effect. He also says that within ten years there may be more robots than people, each person having their own copy, their own avatar.

— Zhang Ya-Qin
26:15

Large models relieved all three long-standing problems of autonomous driving at once

Zhang Ya-Qin reduces the difficulties of autonomous driving over the past decade-plus to three: not enough test data; too many corner cases, safety scenarios you can't encounter; and technological fragmentation — one model for maps, one for vision, one for lidar, with perception, fusion, planning and decision-making pieced together block by block, mixed with rules and neural network algorithms. After large models appeared, all three were relieved at the same time: generative models can create data, and simulation gets faster; large models come with common sense, so scenarios never encountered can also be solved; and end-to-end puts all the models into one box, one input and one output, with rules at most as a fallback. He says Musk's end-to-end model is an important product that let everyone see the dawn.

— Zhang Ya-Qin
29:16

In the past a new scenario meant a new algorithm; now everything is a Transformer

Zhang Ya-Qin points out the fundamental difference between this generation and the previous generation of deep learning: in the past you used convolution one moment and RNN the next, different algorithms for different inputs, with the outputs then fused; now no matter what it is, it is the same Transformer, tokenized, token based. This uniformity changes the nature of the robotics problem too. Robots used to use reinforcement learning, one environment and one agent learning a policy, and applied to a physical body it often didn't work — real to sim didn't work. Their own method is called RS2, connecting the real world and the physical world, learning first in digital space and then transferring back to the real world, solving the data shortage. More crucially, large models give robots a brain — in the past a robot could understand your words but not what lay behind them, it had no common sense, no comprehension, no reasoning; now you can direct it to decompose a task like "take the dirty clothes to the dry cleaner" into actions.

— Zhang Ya-Qin
39:17

Biological intelligence in twenty years, but consciousness won't grow out of silicon

Zhang Ya-Qin believes biological intelligence needs longer, achievable within twenty years. On the path, he is more optimistic about non-invasive sensors and brain-computer interfaces, used first for treatment — stimulating neurons in the blind, stimulating the hearing neurons of the deaf, repairing the central nervous system of the disabled, and for Alzheimer's, ADHD and autism — and only afterwards extending to the intelligence of normal people. But he draws a clear line: he does not think artificial intelligence can currently produce consciousness, nor does he believe there will be a new self-awareness, a new soul. The reason is that "we don't even know how we humans produce consciousness," and "we have no way to create something we don't understand." He cites Richard Feynman to support this.

— Zhang Ya-Qin
43:21

The future is a new species, but control stays in human hands

Zhang Ya-Qin's view of species: the future is a new species, only this species becomes extremely intelligent and enormously capable, but it is still an extension of humanity, controlled by humanity. He takes the Upper Cave Man of thirty thousand years ago as a reference — by then humans already had tools and fire, but the tools were very primitive; a hundred years from now, looking at today's phones, internet and PCs will be like us today looking at stone tools from thirty thousand years ago. He especially stresses the change in the speed of evolution: from hunting to agricultural society, two or three thousand years brought little change; the real great change was the three hundred years after the Industrial Revolution, and after the steam engine appeared human evolution was no longer Darwinian natural evolution but non-linear, exponential evolution. So he says that in a hundred years, maybe thirty, today's humans will look very simple.

— Zhang Ya-Qin
47:22

In thirty years, driving will be like taking a carriage in New York today

Zhang Ya-Qin's concrete imagination of ten years from now: many cars are driverless, many robots enter the home, like today's refrigerators and televisions — when you're not home they check whether the windows are closed, whether anyone has broken in. Very few people will drive, many will choose not to, and the money for buying a car may be spent on buying a robot. Thirty years from now, seeing a person drive will be as novel as seeing a horse carriage in New York today; a person driving may need special approval, a special license. He uses elevators as an analogy: thirty years ago there was still a person sitting in the elevator pressing the floor for you, now almost none, but some places in Britain still keep manual elevators, and it has instead become a kind of high-end service.

— Zhang Ya-Qin
48:22

Longer lifespans will first change the work week, then population structure

Zhang Ya-Qin judges that in thirty years human lifespans will increase substantially — perhaps not immortality, but a hundred years will become the norm, with some reaching a hundred and fifty. The social consequences he extrapolates: working hours keep shrinking — after the Industrial Revolution people worked seven days a week, later six; when he went abroad it was still six, and when he came back it was five; now some places in Europe are already at four, and in the future people may work one day a week and spend the rest of the time doing what they like. He stresses this is not an unemployment narrative but a matter of social productivity rising substantially and social structure changing: human lifespans lengthen, the natural birth rate falls, fewer people are born, and the population of developed countries is already shrinking. He also says that in the future an eighty-year-old may still be a youth, in his prime.

— Zhang Ya-Qin

In their own words · checked verbatim

He said, when you are not worrying about this question, when you are not talking about this question, you have become first-rate.

他说当你们在不为这个问题的时候,不讲这个问题的时候,你变一流了

Zhang Ya-Qin4:04

No matter how many books you read, you still can't swim — you still have to go swim.

你读再多书,怎么游泳,你还是不会游,你还要再游泳

Zhang Ya-Qin13:08

A robot is an open problem; autonomous driving is a close problem.

机器人,他是个open problem,自动驾驶,他是个close problem

Zhang Ya-Qin14:08

It can understand your words but not what lies behind them — it has no common sense, no comprehension.

他听得懂你那字,他不知道后面什么意思,没有这个常识,没有理解力

Zhang Ya-Qin31:16

Thirty years from now, seeing someone drive will be like seeing a horse carriage now — it becomes something novel.

三十年之后可能,你看到有人开车,就像现在以后,开马车一样,变成新型的东西了

Zhang Ya-Qin46:22

Figures

Judged time for information intelligence to reach human levelwithin five years11:08
Time needed for humanoid robots (high-level physical intelligence)ten years14:08
Time for biological intelligence to be achievedwithin twenty years39:17
Common back-end proportion across the three robot types70% to 80%17:14
Share of time spent recruiting in Microsoft Research's early days80%5:04
Expected normal lifespan thirty years from now100 years, some reaching 15048:22
Projected future work weekwork one day a week48:22

Glossary

RS2 / real-to-sim then back to real
Zhang Ya-Qin's team's method: convert real scenarios into the digital world for training, then bring them back to the real world for deployment.
corner case
Rare scenarios in autonomous driving testing that are hard to encounter yet determine safety.
end-to-end
Combining modules such as perception and planning into a single model, with one input directly producing one output.
world model
A model that constructs and simulates the physical world in digital space, used to train embodied intelligence.

How to listen

Who it's for

Founders and investors watching the AGI roadmap and the pace at which autonomous driving and embodied intelligence land; anyone who wants a timetable and classification framework from an academician.

Skip

The book recommendations after 51:27 can be skipped; they are unrelated to the main argument.