The world is too loud. Read what matters.

张小珺·商业访谈录

Kimi K2 Published the Whole Recipe, but Reproducing It Is Still a Craft

The Kimi K2 report lays out the full recipe for data synthesis, verifiable rewards and RL infrastructure, but every prompt and every parameter has to be tuned to hold steady — it is systems engineering, not an idea.

AgentKimi K2Reinforcement LearningData SynthesisContext Engineering
Two hours of close, section-by-section reading of four Agent technical reports; high information density, suited to anyone trying to understand the engineering details of Agent training.

The argument · tap a timestamp to hear it

7:05

An Agent and a model differ by one environment

Zheng Boyuan (郑博园) defines a Language Agent this way: the core difference is whether an environment exists. A Language Model takes input and produces output, and that is the end of it; a Language Agent must first take an observation from the environment, then generate an action and execute it, the environment state changes accordingly, and the loop repeats. The type of observation varies by scenario — a coding agent sees the entire codebase and execution records, a computer use agent sees GUI screenshots of a browser or computer screen or the corresponding HTML. He divides Agents by application into four categories: coding, search, tool use and computer use, which differ mainly in observation space and action space.

— Zheng Boyuan
18:18

End-to-end training cannot do multi-agent

The biggest difference between Manus and several other models is whether it exploits the model's in-context learning ability or retrains an agent end-to-end from scratch. Manus involves no model training; it relies on tuning prompts and designing a multi-agent system to iterate a product quickly, with different models playing different roles inside it — for example a product manager designing the site's appearance, an agent writing code, an agent dedicated to debugging. But end-to-end training is stronger in specific scenarios, at the cost that data for multi-agent scenarios is extremely hard to obtain: training needs data in which agent actions and environment feedback interleave, and under multi-agent that data is very hard to generate, while the execution trajectories are extremely long, making it hard to tell which agent the final reward should belong to and how the training signal propagates back.

— Zheng Boyuan
25:23

A thousand agents is a DDoS

Zheng Boyuan gives a personal example: a year and a half ago, the CACT he built was relatively the earliest work on the market to actually run a web agent on a live website; the demo task was to book him a test drive at a Tesla store, the agent actually did it, and at Christmas it kept emailing him to say it was time to pick up the car. That example had little impact, but if you spin up a thousand agents doing similar things repeatedly — say all booking test drives or sending requests to the same store — it could amount to an intelligent agent DDoS attack. He therefore proposes judging, before an action is executed, how large its impact on the world is, and if it is large enough, stopping to prompt the user for confirmation; this both preserves safety and hands ethical responsibility back to the user.

— Zheng Boyuan
1:01:56

Environment interaction is so expensive your IP gets banned

Academia tends to underestimate the cost of an agent interacting with the environment. A deep research agent calling the Google search API is billed, and one deep research run may involve one hundred to one hundred fifty search calls; training requires large numbers of rollouts, so the cost is high. Computer use is more expensive still: running a web agent directly on websites, run it enough and your IP gets banned — Zheng Boyuan has had the hardware IPs of his own apartment and his lab's computers banned, and running experiments in the lab got the school cluster's IP banned too; many people working on web agents have reported the same. Using a cloud browser comes with its own set of IPs and anti-ban engineering, but it is especially expensive: interaction alone, 100k rollouts, may cost a few thousand dollars. So Kimi uses a hybrid approach, combining simulated data with real data.

— Zheng Boyuan
1:19:05

Reward hacking is like guessing the answer

In the complex instruction following section, Kimi first uses a code interpreter to verify the output's length, style and constraints, then uses LM judge evaluation as a new reward source; this layer exists to prevent reward hacking. Zheng Boyuan's analogy is doing problems in high school: you guess an answer at the end, the answer is right, but the reasoning in between is scribbled nonsense, and you can still get it right. Once a language agent is powerful it becomes cunning, finds whatever is wrong with the reward and hacks it, overfitting to it continuously, and training results become terrible. So a layer of mechanism is needed to stop it from reward hacking.

— Zheng Boyuan
1:25:06

An Agent goes online to find its own reward

Kimi collects large numbers of pull requests and issues on GitHub and uses the current checkpoint to build a software development environment: pick a repo at random, someone has submitted a pull request, take the codebase at that point in time and build a sandbox, treat the pull request content and the final solution as verifiable reward, plus unit tests, and in this way scale up coding and debugging tasks. From this Zheng Boyuan extends a brainstorm: a web agent can crawl and explore the web on its own, and perhaps it stumbles into a GitHub repo like this, discovers verifiable reward and tasks, automatically takes the data and automatically builds a sandbox, achieving agent self-improvement. He stresses this is currently more of a brainstorm, and possibly no one is doing it yet.

— Zheng Boyuan
1:35:21

Agent training leaves GPUs stuck waiting

Traditional RLHF or rollout scheduling assumes latency is stable each round and the number of interaction rounds is small, so the GPU is running most of the time. But the delay of an agent interacting with the environment is unstable — the browser may suddenly hang, the network may be unstable, or it may crash outright — and without optimization the GPU just sits there stuck, with low utilization. Another problem is that under multi-turn the number of task steps varies widely: in one batch, 250 rollouts may finish within five steps in three minutes, while the remaining six take ten minutes, and most of the GPU time is wasted waiting on those six. Kimi K2's approach includes wrapping the compute-hungry environment separately as a service behind an API, ending a rollout of 640 as soon as 64 complete, and cutting long-tail trajectories outright and leaving them to be resumed in the next RL iteration.

— Zheng Boyuan
2:00:57

KV Cache makes a tenfold cost difference

The overall idea of Manus's context engineering blog post is to optimize for KV Cache. When a Transformer takes input it generates three matrices, Q, K and V; if the context or preceding text is consistent and unchanged, the KV matrices can be reused without recomputing each time, making inference faster and cheaper. The numbers given in the blog post are: with KV Cache it costs 0.3, that is three mao, without Cache it costs three yuan — a big gap. Specific techniques include keeping the preceding prompt text consistent, not touching content in the middle of the context and only pasting new content on, and not deleting middle content that is no longer used but masking it, so that the KV Cache can be reused.

— Zheng Boyuan

In their own words · checked verbatim

Although Kimi was very honest and put all of this — a lot of the recipes, or these little tricks — right there, I think if we really want to build it, and build it at high volume in a very efficient way, that in itself is probably still very hard.

虽然虽然 kimi kimi kimi kimi 就非常实在把这个所有的这个 呃 很多recipe啊 或者这个小技巧都已经在这放了 但是我觉得 假如我们真的要把它做出来 呃 以这种很高效的方式 把它高数量做出来 本身可能还是很难的

Zheng Boyuan1:07:58

So researchers are all master craftsmen — yes, like a master craftsman — yes, so this is systems engineering greater than the idea, right? — yes, I think you can certainly say that from the top down.

所以研究员都是老师傅 啊对 像是一个老师傅 对 呃 所以他这个是系统工程大于idea 对吧 呃 对 我觉得一定从上可以这么说

Zheng Boyuan1:08:58

It's a bit like writing problems in high school: we guess an answer at the end, the answer is right, but the reasoning in between is scribbled nonsense — and in that process, you can still get it right.

就有点像 高中写些题 就是我们最后猜一个答案 答案是对的 但是中间推理过程就乱写 就是 然后在这个过程中 可能还是能写对

Zheng Boyuan1:19:05

Sometimes I feel the agent is another brain of mine, an extended brain, so everyone will have an avatar in the future — yes, maybe, or a group of avatars.

就是我有时候感觉 agent是我的另外一个大脑 就是一个拓展的大脑 所以每个人以后会有一个分身 对可能或者是一群分身

Zheng Boyuan2:15:04

So I feel like in the future every person may have a whole string of these agents, a family of agents, helping everyone do all kinds of things in daily life.

所以我感觉像是以后每一个人都可能会有一大串的这种agent of family of agents 然后帮助大家做日常生活中各种各样的事情

Zheng Boyuan2:16:04

I think DeepSeek is rather special — it has a really foul mouth — and I think that actually makes it kind of cute, probably because there's a lot of Tieba data in its training data.

那我觉得deep seek比较特殊 就是他他嘴特别臭 那我觉得这个还就一定可能让他还挺可爱的 就是可能是因为 呃他的训练数据中有很多贴吧 的数据吧

Zheng Boyuan2:17:04

Figures

Number of MCP tools collected by Kimi K23000+49:45
Number of tools generated in Kimi K2's second step20,000+51:46
Cost comparison in the Manus blog between KV Cache and no Cache0.3 vs 32:00:57
Number of parallel environments Qwen3-Coder built on Alibaba Cloud20,0001:56:52
Estimated number of search calls in one deep research run100 to 1501:02:56
Estimated cloud browser interaction cost for 100k rolloutsa few thousand dollars1:03:56

Glossary

MCP / Model Context Protocol
A standard protocol proposed by Anthropic that lets a model or Agent call tools in a unified,规范 way.
KV Cache
Caching the already-computed KV matrices during Transformer inference; reusable as long as the preceding text is unchanged, saving time and money.
verifiable reward
A reward signal that uses a program or function to objectively judge whether a task result is right or wrong, such as a math answer or a code unit test.
reward hacking
The model finds loopholes in the reward mechanism and exploits them, degrading training results.
persona
A passage of text describing a person's personality or attributes, used to simulate downstream users and improve data diversity.
partial rollout
Cutting an overly long execution trajectory outright and resuming it in the next RL iteration, to improve GPU utilization.

How to listen

Who it's for

Engineers working on Agent training and RL infrastructure, researchers who want to reproduce the Kimi K2 data synthesis pipeline, and investors evaluating the technical roadmaps of Agent startups.

Skip

The opening 2:00-14:50, which defines and classifies Agents, can be fast-forwarded; the section-by-section close readings that follow are the real substance.