The world is too loud. Read what matters.

乱翻书

The deciding factor for a Personal Agent isn't the model — it's memory and proactivity

The Today team believes models are approaching AGI and costs are falling tenfold every year, so the bottleneck has shifted to memory and proactivity; and a product that does proactivity badly just becomes another spammy push notification.

Personal AgentMemory SystemsProactivityProduct DesignAgent Frameworks

The video won't play here. Listen to the audio instead:

Two founders get very concrete about the engineering details of a Personal Agent: how memory is stored, how proactivity is evaluated, why intent routing is a disaster. High information density, but the product is not yet validated.

The argument · tap a timestamp to hear it

14:15

The model is no longer the bottleneck — cost is

Gao Ce previously worked on model infrastructure and vector databases, always writing code by hand, until a model release this February made him feel intelligence had taken a leap, to the point where he thinks coding has reached AGI level. His reasoning: if coding can still improve tenfold or eightfold every year, and cost falls tenfold every year following scaling law, then in a year or two there will be cheap intelligence on your phone that is stronger than an average person. So what a Personal Agent lacks is not model capability, but how to use it. This judgment determines that Today's product focus is not the model, but the harness and product design.

— Gao Ce
20:16

Serving people with lots of digital assets first is a tradeoff you have to make

Gao Ce explains why early users are always extremely busy people willing to tinker: for a Personal Agent to create value, the precondition is that you have left enough context in the cyber world. For someone who doesn't use email and has no online footprint, the Agent cannot quickly learn your profile and can only have you tell it bit by bit, making the onboarding barrier extremely high. That's why OpenClaw first took off among programmers, because they are at a computer every day and leave behind a large amount of context. A product serving only high-net-worth individuals cannot scale, but before the problem of unifying physical-world and online context is solved, you can only start here.

— Gao Ce
26:27

The opposite of Personal AI isn't office work — it's high delegation cost

Suki pushes back on the question ‘you're doing a personal agent, why are you also doing office work’: the opposite of Personal AI is not work, nor an office agent, but high delegation cost. In the past, delegating something to an agent meant giving it a lot of context, whereas this generation of products goes directly into your life, connects to your apps, knows your colleagues and relationship network, and often can get something done in a single sentence. Gao Ce adds that for the vast majority of people, especially Chinese people, life and work are hard to separate — the busier you are, the blurrier the boundary — so generalizing from life to work is easier than the reverse.

— Suki
29:29

Big companies are good at big resource plays, not at seeing and persisting

Suki says that as far as she knows, Qwen and Doubao saw the OpenClaw direction mid-year and explored it for a while, then quickly gave up. Her learning is that what big companies are best at is not seeing something and believing in it, but big resource plays and burning money. Right now this direction has no PMF and a very small TAM — even at 15% paid conversion, the user volume would still look like a small market to a big company. Last year DingTalk was 4 billion, Feishu 3 billion, WeCom around 3 billion — under 10 billion combined. So big companies usually wait until the TAM is big enough before investing regardless of cost.

— Suki
35:33

Raising a bunch of Agents is emotional value, not efficiency

Last December Gao Ce himself raised three agents — a design chief, a product chief, and a project chief — and in the end only wanted to chat with the project chief, because the context was all in one place and he didn't need to explain it repeatedly to three roles. He thinks multi-agent may be a false need on the consumer side; many people raising a bunch of crayfish are mostly after the emotional value of ‘I have a bunch of underlings doing work’, but in actual use it wastes tokens — coordination between different roles also costs your tokens and money. The real value of multi-agent is two brains supervising each other, such as in evo and testing scenarios, but most personal agents' daily use doesn't need that.

— Gao Ce
44:38

The outbox has higher information density than the inbox

Suki describes a concrete rubric for memory construction: after Town gets email authorization, the naive approach is to scan the inbox and outbox for clues, but the core observation is that the emails you send carry higher information density, because sending is done with intent, and you describe your intent clearly — for example, when you want to collaborate with someone you introduce yourself first. The inbox is other people's intent toward you; the outbox is your intent toward others. Email is a protocol that dates back to the 1960s, with a large amount of protocol and side-channel signal behind it, and whether you can capture that signal in the product is the difference in memory quality.

— Suki
47:38

Proactivity is an open scenario — you can't run AB tests on it

Gao Ce contrasts work scenarios with proactivity: work scenarios are closed problems, like the benchmark Workbody open-sourced — give it a workspace, finish the task, and you're done, easy to evaluate right or wrong. But proactive is an open scenario that uses a large amount of your incremental information, with goals that are highly subjective, personalized, and multi-dimensional, related to push frequency, quality, and even your mood that day. In the early stage a product cannot run GSP or AB tests to validate proactivity, and it will be very difficult in the future too, so it heavily tests the product lead's understanding of the world and information architecture — it's something engineered and product co-designed.

— Gao Ce
1:02:50

Stacking an intent model in front is an engineering disaster

Gao Ce on intent recognition: the traditional approach is to stack a small intent model in front, then dispatch to different large models by intent — for example, recognizing a weather question and using a very small model. But the drawback is that intent maintenance gets harder and harder; as the product horizontally expands across intents, you may end up stacking thousands of intents, and products with strong generalization like Doubao, Yuanbao, and Workbody are especially hard to maintain. Moreover, once you stack an intent model, the context seen below is fragmented, and how messages are shared across multiple sub-models is a big problem — for engineering it's quite a disaster. Manus published a blog in July 2025 about a context-aware state machine, but didn't go into implementation details.

— Gao Ce

In their own words · checked verbatim

The opposite of personal AI is not work, not an office agent — the opposite of personal AI is high delegation cost.

personal AI 他的对立面并不是工作 并不是办公agent personal AI的对立面是高委托成本

Suki26:27

I hear a lot of people raising a bunch of crayfish, and really it's more about the emotional value of seeing them work — that I have a bunch of underlings doing work for me. But if you actually use it, you'll find it's wasting your tokens.

我听到很多人 养一堆小龙虾 其实更多的是看到 他们干活的 一种情绪价值的感受 就是我有一堆小弟 给我干活 但实际上 如果真的用起来 你会发现 它其实是在浪费你的token

Gao Ce38:33

My feeling now is that model generalization isn't as strong as imagined — the value of the product manager is very large. And models are actually becoming more and more homogeneous. Once models become infrastructure in the future, the value of the product, or the value of the product manager, I think will only grow.

我现在的感受是说模型泛化性没有想象中那么强 就是产品经理的价值是非常大的 以及模型其实现在它的同质化其实是越来越重的 就模型未来成为基建以后 产品的价值或者说产品经理的价值 我觉得是越来越大的

Gao Ce58:46

I think in the short term everyone is indeed overestimating this field, but in the long term it may be underestimated. In the short term, I think it's very overestimated.

我觉得是短时间内 大家确实是高估这个领域的 但是长时间 他可能是被低估的 但短时间我觉得是非常被高估的

Gao Ce1:22:04

Figures

Today's timeline from idea to launchIdea in October last year, team formed in February, internal testing started in April2:02
Annual revenue of DingTalk, Feishu, WeComDingTalk 4 billion, Feishu 3 billion, WeCom around 3 billion — under 10 billion combined30:30
Suki's unread email countOver four thousand unread emails57:45
Model training data mixXiaomi's RL training process was open, with a data mix of 60-70% coding and 3.5% chat1:17:02

Glossary

Personal Agent
An AI assistant that remembers you and proactively does things for you, as opposed to a tool that requires you to initiate a task.
harness
The framework built around a model for tool calling, context management, and task execution.
proactive
An Agent's ability to push information or execute tasks without waiting for you to ask.
KVCache
A mechanism in large-model multi-turn dialogue that reuses previously computed results to lower cost.
rubrics
Human-designed judgment criteria used to guide a model toward more accurate judgments in specific scenarios.

How to listen

Who it's for

Founders and product managers building AI applications and Agent products, especially anyone currently wrestling with memory architecture, proactivity design, and intent routing schemes.

Skip

The closing reflections and goodbyes after 1:36:16 — low information density.