The world is too loud. Read what matters.

张小珺·商业访谈录

The bottleneck for agents isn't intelligence, it's that they can't learn to become experts

General intelligence is already good enough and cheap; what's truly scarce is specialized intelligence — letting an agent, like an intern turning into a seasoned veteran, continuously learn the world model of some small world.

AgentContinual LearningWorld ModelStartupsHCI
The technical history is clearly laid out, and the second half has the highest density of argument on continual learning, world models, and why the GUI won't disappear — suited to engineers and founders who want to figure out where agents go next.

The argument · tap a timestamp to hear it

6:08

Logical agents died of the knowledge acquisition bottleneck

Su Yu calls the expert systems of the 1950s to 1990s logical agents: write a domain expert's knowledge into a logical language, pair it with an inference engine, and when a new problem arrives, do logical deduction. Its memory is just a finite set of logical statements, and its expressive power is bound by the logical language itself — the vast majority of things in the world can't be expressed in simple logic. Its autonomy is also reduced to ‘take a problem, deduce an answer’. What really crushed it was the knowledge acquisition bottleneck — relying on engineers to interview experts and then translate that into logical form, a process that was painful, inefficient, and of limited effect, directly causing the AI winter of the 80s and 90s.

— Su Yu
13:17

Deep RL agents get only one forward pass

After 2000, neural agents produced representatives like AlphaGo in deep reinforcement learning, but from the perspectives of memory and autonomy they remain very limited: the subject is just a very small neural network, tens of millions up to at most 100 million parameters, playing only one game or one class of games, with frames as input and actions as output. Its reasoning is implicit, hidden inside a single forward pass — no matter how complex the situation, the compute it can use is one forward pass. Humans aren't like that — the more complex the situation, the more compute reasoning clearly takes. Another hard flaw is extremely poor sample efficiency; a simple game might take millions of rounds to learn.

— Su Yu
21:27

Language is the scaffold for this generation of agents

Su Yu, Yang Di, Yao Shunyu, and Yu Tao specifically made a tutorial in 2024 to define language agents. This generation of agents is based on LLMs, and the biggest difference is that language can be used as a scaffold: perception uses language understanding, and the forms of interaction are far more flexible; reasoning uses chain of thoughts, and the more complex the task, the more tokens are produced, each token being a forward pass and a certain amount of compute, thus achieving adaptive computing. Language is also the medium for taking action, including formal language and machine language, and can do all sorts of things in the digital world. From the memory perspective, large model training itself is a process of using language as a scaffold and forming a representation of the world through compression.

— Su Yu
48:56

The OpenClaw moment and the ChatGPT moment are the same kind of moment

Su Yu thinks the two are highly similar: the underlying technology was actually ready long ago. Before ChatGPT, LMs had already developed for several years, from BERT, ELMo, GPT-1 to GPT-3; ChatGPT just finetuned the model to be more like a chatterbot and released it directly to the public — the underlying technology didn't change much, what changed was the interaction form, and that change became the fuse. OpenClaw is the same: most people doing agents who look at its code base will have a feeling of nothing is new here, but it is also a profound change in interaction form: you can interact within instant messaging software, it has an independent environment, it's always on 24 hours a day, and it's open source, ignoring permission and safety, with everything opened up. He believes that looking back in two years, OpenClaw's influence may be of a similar scale.

— Su Yu
57:23

Foundation model intelligence has crossed the threshold; what's missing is people to capture the value

Su Yu cites former Google CEO Eric Schmidt's observation: the US is generally much slower at the application layer, while China has always moved fast on front-end technology applications, and in the AI era this is a big advantage. The reason is that the intelligence of foundation models has already crossed the threshold, and for many useful things it's good enough. Many things weren't done before because the friction was too high and the economics didn't add up; now AI capabilities can greatly reduce that friction, and many things go from not worth doing to worth doing, which creates commercial value. What's missing is people with enough insight and execution to discover and capture that value. He admits the process will involve waste — for example, first spending money to install OpenClaw, finding it useless, then spending money to hire someone to uninstall it — but for society as a whole it's still a positive development.

— Su Yu
1:03:26

Once general intelligence becomes cheap, differentiation comes from specialization

Su Yu says AI has now reached a stage where general intelligence is very strong: in the digital world, give Cloud Code or OpenClaw a problem at random, and as long as it isn't highly specialized and the necessary information is there, there's roughly a 60 to 70 percent chance it gets it right. So what's missing is specialized intelligence: when general intelligence becomes standard equipment, differentiation comes from specialization. The world isn't one whole world; it's made up of millions of small worlds — every profession, every domain, every company, every piece of software, every website is a small world, and the entropy of these worlds adds up to almost infinity. It's impossible for a single agent or a single model to capture all of it; there will necessarily be a process of adaptation and specialization.

— Su Yu
1:14:48

A world model isn't just video prediction; it's an intern becoming an expert

Su Yu's definition of world model is broader than most people's. When people mention world model they usually mean vision-based models doing next frame prediction or 3D reconstruction; he thinks this is important and is also a capability LMs lack, but world model is the most important concept in the whole of human intelligence. He gives an example: a fresh college graduate joins a company as an intern, on the first day has no idea what the work is, but can continuously learn through learning on the job — what the company's org chart looks like on the surface versus in reality, who actually has a say, who to go to for approval, how the software is used, what the workflow is, the theory of mind between people — all of this is part of the world model. The model of this microworld is obviously not a video model, and many parts are naturally symbolic.

— Su Yu
1:44:05

CLI won't fully replace GUI

Su Yu thinks the GUI won't disappear: humans are visual animals, the brain is just wired that way, and HCI research shows that the same thing visualized is understood by the brain a fraction of a second faster; the GUI also has practical benefits for validation, winning trust, and auditing. Whether an agent needs a GUI is another matter, but the GUI is the de facto interface of the entire digital world — 99% of things already have a GUI to interact with, and the GUI has already encoded a great deal of knowledge, constraint, and business logic in its design process. If an agent uses the GUI well, it can piggy back on all of this accumulated knowledge, rather than reinventing a set of CLI or API wheels, and only in this way can it immediately reach all corners of human society, especially long tail scenarios. He also gives the example of the semantic web: Tim Berners-Lee pushed it for over twenty years, and adoption is still very low, because society doesn't work that way.

— Su Yu
2:06:19

The real risk is job displacement outpacing new job creation

Su Yu doesn't see the possibility of AI rapidly self-iterating and wiping out humanity in the foreseeable future, because that isn't just an intelligence problem but a lack of higher-level capabilities like innate goals, intention, and survival pressure — right now all purposes are assigned by humans. His real concern is job displacement: if AI agents replace knowledge workers on a large scale, and on one hand can't create enough new jobs to carry the displaced workforce, and on the other hand there's no good mechanism for redistributing the gains to provide a social safety net, with most of the gains captured by a few leading companies or capital, it will have an enormous impact on society. He believes every AI researcher has a responsibility, and that the important thing he can do is democratize access to frontier agent capabilities.

— Su Yu

In their own words · checked verbatim

But at the end of the day, all these things are rapidly converging, and at the end of the day what everyone wants is a Universal Digital Agent.

但At the end of the day 就是最后这些东西都是在快速的converge 最后At the end of the day 大家想要的就是一个Universal Digital Agent

Su Yu41:09

Most people doing agents who look at OpenClaw's code base will probably have a feeling of nothing is new here, that there's no innovation in this place, but in fact it is also a profound change in interaction form.

就大部分做agent的人去看open cloud的这个code base的话 可能会有一种nothing is new here 就这地方没有什么创新的这种感觉 但实际上它是一个也是一个交互形式的一个深刻的一个变化

Su Yu51:15

This model is obviously not a video model, but vision is of course a very important part of it, yet there are obviously more parts that are naturally symbolic.

这个model它显然不是一个video model 但vision当然是里面很重要的一部分 但显然也有更多的部分 它是天然就是符号化的 symbolic的

Su Yu1:16:50

This kind of individual thought doesn't need a language, but civilization needs a language.

这种individual thought doesn't need a language 但是civilization needs a language

Su Yu1:36:01

But now, for example, I first put out a standard called MCP, or I put out a standard that makes everyone go write CLI, and then you expect all industries to adopt this in the next few years — that's almost impossible.

但你现在 比如说我就先出来一个标准叫mcp 或者我出来一个标准 就是让大家都农进去写cli 然后你指望所有的行业都 在未来几年去adopt这个事情 这是几乎不可能的

Su Yu1:43:05

Because he hasn't learned it, even if he's done it, he has no effective way to learn it like a human and make it a part of my expertise, which is why it leads to this instability.

因为他没有学过 他即使做过 他也没有一个有效的方式 把他给像人一样去学会 成为一个part of my expertise 所以他才会导致这些不稳定

Su Yu1:52:13

That is, all the way continue learning, all the way world modeling.

那就是 All the way continue learning All the way word modeling

Su Yu2:15:23

Figures

AutoGPT GitHub Starsabout 180,00034:03
Neocortex share of the brainabout 70%1:21:53
Number of cortical columns in the human brainabout 150,0001:24:54
OpenAI and Anthropic share of market fundingabout 30% to 50%1:07:38

Glossary

semantic parsing
Turning what a person says into a formal meaning representation a machine can read, such as a knowledge graph, database, or executable form for a website.
language agent
An agent that uses language as a scaffold for perception, reasoning, and action, as defined by Su Yu and others in a 2024 tutorial.
world model
An agent's internal representation of the small world it inhabits; Su Yu argues it isn't limited to video prediction but also includes org charts, workflows, theory of mind, and more.
continual learning
A model continuing to learn new tasks or new environments after deployment without forgetting existing capabilities; Su Yu argues its learning target should be the world model.
cortical column
A repeating unit structure in the neocortex, of which the human brain has about 150,000; Jeff Hawkins' theory holds that each cortical column is learning a world model.
non-parametric learning
Giving an agent new capabilities without changing model parameters, by writing MD files or harnesses and the like; OpenClaw's skills fall into this category.

How to listen

Who it's for

Engineers, founders, and investors who want to understand the arc of agent technology and its next bottleneck, especially those concerned with continual learning, world models, and agent product form.

Skip

The rapid-fire Q&A after 2:10 can be skipped, and the self-introduction around 2:00 can be fast-forwarded.