The Second Half of AI: The General Method Exists, the Hard Part Is Finding the Task
Yao Shunyu says the main thread of AI has moved from "building weapons" to "choosing the battlefield" — the method is finally general, but the real bottleneck isn't reasoning ability, it's that models can't get the context inside people's heads.
The video won't play here. Listen to the audio instead:
The argument · tap a timestamp to hear it
From vertical thinking back to general thinking
Yao Shunyu tells the history of AI as a fragmented curve: starting from Newell and Simon in the 1960s wanting to directly build an agent, and because it was too hard, the field kept cutting the problem smaller — first vision, then language, then translation, becoming more vertical with each cut. He believes the significance of scaling laws and a batch of research breakthroughs after 2015 is that they reversed this: from subdivided thinking back to the thinking of building a more general system. This is also the underlying coordinate for judging what he should do.
— Yao ShunyuThe essential difference of language agents is the ability to reason
He gives a very concrete contrast: AlphaGo can only play Go, and switching games requires retraining; but a human can pick up a new game on the first try. His explanation is that humans think — seeing a light is off, inferring there might be danger, then combining context to judge the light is behind them, so they walk backward first. The sufficiently strong prior provided by language models makes reasoning possible, and reasoning can generalize across different environments. So the essential difference between language agents and all previous agents is not the interface, but the generalization brought by reasoning ability.
— Yao ShunyuReward should be outcome-based and rule-based
For any RL task, he insists on two things: reward is based on outcome rather than process, and it is white-box and clearly computable, not based on human or model preferences. The reason is that as long as it is based on process or preferences, hacking will occur — you might get a very elegant piece of code, but it doesn't solve the problem. He gives math as an example: if the answer is 3, it's 3; if it's not 3, it's wrong, no room for interpretation. He says when doing webshop, the hardest part is not setting up the environment, but designing a task that is both difficult and genuinely valuable, with reward that isn't noisy.
— Yao ShunyuCustomer service needs pass^k, not pass@k
The traditional metric in coding is pass@k: the probability of succeeding at least once when running the same task k times, so some report pass@100. But tasks like customer service need the exact opposite metric, which they defined in a paper last year as pass^k — the probability of succeeding all k times. He judges that people currently focus more on success rate or pass@100 because they are still doing benchmarks rather than real applications; once you accept this shift, you'll realize some applications must optimize for robustness, and there's still a lot of room for progress there.
— Yao ShunyuThe opportunity for startups is to design new interaction paradigms
He flips it around: what startups should worry about is not model capability overflowing, but model capability no longer overflowing — that's when you really can't do anything. His framework is: model companies' products are all ChatGPT-style, human-like interactions; the biggest opportunity for startups is to use the model's general capabilities to create different interaction paradigms. Cursor is an example — humans don't interact with each other that way. He adds a key constraint: doing new interaction paradigms and the model continuously overflowing are both indispensable; if your interaction paradigm is very similar to ChatGPT, then what reason do you have not to be replaced by ChatGPT?
— Yao ShunyuWhat models lack is not intelligence, but context
This is his core answer to "why do models with such strong reasoning and exam performance not create enough economic value": models lack context. A huge amount of context in human society exists only in people's brains, maintained in a distributed way — for example, your boss's behavioral habits are hard to summarize in words. So a person who graduated from a second-tier university and is worse at math than O3 can, after seven days at the company, accumulate enough context to do better than O3. He believes if this problem is solved, the utility problem will be largely solved.
— Yao ShunyuThe environment is the outermost layer of the memory hierarchy
He quotes a line from von Neumann's last book before he died, The Brain and the Computer: environment is always the most outer part of the memory hierarchy. Analogous to a computer going from CPU cache to memory to hard drive, the outermost layer is always the external world — plugging in a USB drive, uploading things to the internet, burning a CD. For humans, the outermost long-term memory is your notebook, Google Doc, Notion. He thus places MCP under the same logic: essentially a way to hack your context.
— Yao ShunyuTasks that can be defined as exams are not far from being solved
He proposes what he thinks is the most important basic assumption to overturn: currently evaluating something is based on 500 tasks, each run 500 times, summing parallel data into reward. But what matters for people at work is how much better they get after 30 days or a year, not how well you do on your first day at the company in 100 parallel universes. His inference is: once you can define an exam or a game, it's not far from being solved; the real world is hard precisely because it has no pre-designed reward and standard answer.
— Yao ShunyuIn their own words · checked verbatim
Before, it was like I had many monsters, so I needed to build all kinds of weapons for different monsters to fight them. Now I have a general weapon, I have a machine gun, so the question I need to think about now is where to point the gun.
我们之前是就有点像我有很多怪兽 那我需要去为了不同怪兽去造各种各样的武器 去来打这些怪兽 现在我有一个通用的武器了 就我有把机关枪 那现在我要思考的问题是我要朝哪里去开枪
Yao Shunyu38:36
Actually, I think the hardest part of doing any RL task is how to define the reward.
实际上我认为做任何的RL task 最难的部分其实是怎么定义reward
Yao Shunyu40:40
If your interaction paradigm is very similar to ChatGPT, then what reason do you have not to be replaced by ChatGPT?
如果你的交互方式很像chatGPT 那你有什么理由 不被chatGPT取代
Yao Shunyu49:44
Although you are not as smart as O3, you have this context, so you do better than O3.
虽然你没有O3聪明 但是你有这些context 所以你做的比O3好
Yao Shunyu1:00:51
Essentially environment is always the most outer part of the memory hierarchy.
Essentially environment is always the most outer part of the memory hierarchy
Yao Shunyu1:42:18
Because we found that once you can define an exam or a game, it's not far from being solved.
因为我们发现 一旦你可以定义考试或者一个游戏 那离它被解决也不远了
Yao Shunyu1:45:18
If someone else can do it it's okay to let them do it.
If someone else can do it it's okay to let them do it
Yao Shunyu2:25:53
Figures
| Years Yao Shunyu has done agent research | 6 years | 7:06 |
| Metric often reported for coding tasks | pass@100 | 46:44 |
| Multiple of token consumption per user for agents relative to chatbots | 500 to 1000 times | 1:23:01 |
| Time Yao Shunyu spent rewriting The Second Half blog post | about 2 hours | 2:10:37 |
| Yao Shunyu's competition achievement | National silver medal | 2:20:48 |
Glossary
- pass@k
- The probability of succeeding at least once when running the same task k times; a common metric in coding.
- pass^k
- The probability of succeeding all k times when running the same task k times; measures reliability, needed for customer service tasks.
- affordance
- The actionable interface an environment provides to an actor; Yao Shunyu says code is the most important affordance for AI.
- memory hierarchy
- The layered storage structure from cache to memory to hard drive; Yao Shunyu places the external environment at the outermost layer.
- intrinsic reward
- Feedback a system gives itself when there is no external incentive; Yao Shunyu considers this the core mechanism of innovators.
How to listen
Founders building agent applications, investors watching AI product forms, and engineers who want to know "what opportunities remain after model capability overflows".
2:58–17:50 covers personal growth and research history; you can just listen to the conclusions.