Manus's Last Interview Before the Sale: AI Has No Wild West Because There's No New Platform
The mobile internet's wild west came from a hardware generational shift: desktop to phone made giants and individual developers equals. AI's technological leap is bigger, yet no wholly new platform appeared, so giants, startups and individual developers all moved at the same speed. The wild west does not exist.
The video won't play here. Listen to the audio instead:
The argument · tap a timestamp to hear it
AI has no wild west because there is no new platform
Ji Chao attributes the mobile internet's opportunity to a change of hardware medium: from desktop to phone, BAT and the big overseas firms were equals with individual developers, everyone was experimenting, so there was a wild west period. He argues that although AI's technological breakthrough is larger, no wholly new platform emerged, and as a result giants, startups and individual developers reacted at the same speed and moved just as decisively — the wild west does not exist. This judgment directly explains why he later chose not to build models and only to build applications.
— Ji ChaoGPT-3 made him decide on the spot to sell the company
Maggie's team ran from late 2014 to 2018, building two of its own roughly 0.3B models from pretraining up, and solving 16K long context itself. After getting GPT-3 early access, he wrote a casual prompt and found it was fifty-fifty with their own end-to-end model. He realised: it is expensive now, but it is a general solution. Previously the NLP world was sharply divided between people doing information extraction, machine translation and customer-service systems; GPT-3 crushed that line of thinking. His first reaction was to sell the company fast.
— Ji ChaoSearch engines are not a technical problem, they are a data-source problem
Maggie built everything from crawler to index in-house, using no third-party search; he calls it the peak of his engineering ability. But the conclusion was: he is not bullish on building another new search engine, because many data sources have already formed a mutually beneficial, recyclable relationship with Google, and a disruptor cannot repeat what Google spent twenty years accumulating on data sources. From this he sums up the two reasons products fail — technical and non-technical — and the line ‘one step early makes you a pioneer, ten steps early makes you a corpse.’
— Ji ChaoA browser extension is a superb observation window
His reason for choosing Monica was not the product form but unbiased observation: it does not change user habits, users still use Gmail and watch YouTube, their trajectory is not forcibly altered. More crucially, an extension distributes features by context — video understanding only appears when watching YouTube, rewriting only appears in Gmail or Google Docs — dissolving the complexity explosion that comes from stacking features. He quotes the GitHub line: Everything added dilutes Everything else. An extension is not even a product form; it is an empty container, an empty canvas.
— Ji ChaoShipping an uncool product drags you into a self-justification loop
After abandoning the AI browser he gave this judgment: if you finish a product and you yourself don't think it's cool, don't ship it — you should be the person who likes it most, and if you don't like it, how can you expect users to. The more practical reason is opportunity cost: a responsible team that ships a product has to keep maintaining it, enters a self-justification loop, and misses new opportunities that are plainly more valuable. But he leaves room for disagreement: if another team has already taken the browser to a level they are satisfied with, entering the self-justification loop is right, and along that path a completely different product form may grow.
— Ji ChaoCursor exposed that programming is a general capability
The observation that turned the team toward Manus was this: operations colleagues at the company used Cursor to write blogs, data-analysis colleagues used it for analysis and visualisation. Cursor was originally the most professional IDE form, yet a mass of non-target users poured in. They stood behind colleagues and watched: code on the left, a window for chatting with AI on the right, and people who cannot write code used conversation to make AI complete non-coding tasks with code as the medium. The conclusion: programming is not a vertical capability, it is a medium for solving general tasks; Cursor's form is not optimal for them — it should run in the cloud, code should not be the main presentation, and it should target prosumers rather than professional engineers.
— Ji ChaoManus was finished in January and deliberately held until March
Manus started at the end of September 2024 and was basically finished by mid-January 2025, but Ji Chao decided not to ship immediately. The best model available then was Claude 3.5 Sonnet V2, which had preliminary agentic ability but lacked true reasoning. He heard a rumour that a model release would come two months later, so he chose to spend an extra month and a half polishing the product and align the launch with the next model iteration, so that the launch would enjoy the biggest generational jump at the instant it shipped and eat the model spillover.
— Ji ChaoMonica's $12M ARR let them gamble boldly
The browser did not burn much money, and the headcount invested was only a dozen or so people. At the time Monica already had close to 12 million, i.e. $12 million, in ARR, and was a profitable product. He stresses the definition: ARR is MRR times 12, you cannot count annual payments received within a month into that month, and internally they only look at Stripe and mobile MRR data, otherwise there are far too many ways to inflate the number, and that is lying to yourself. Precisely because they had a positive-cash-flow product could they be both bold and rational when building a second curve.
— Ji ChaoAn agent's input/output ratio is 100:1 to 1000:1
In the chatbot era, when doing cost estimates, the input-to-output token ratio was generally calculated at 3:1 — because in a transformer input is prefilling, parallelisable and compute bound, while output is decoding, bandwidth bound, and the pricing differs a lot. But in an agent like Manus, the input-to-output ratio is 100:1 to 1000:1, depending on task type. With output length roughly the same, input length differs by tens or even hundreds of times. This explains why in early 2025 nobody predicted token consumption would grow exponentially — the input/output ratio had changed.
— Zhang TaoContext above 200k is already unimportant
He offered a hot take: more important than longer context is giving the model compaction awareness — awareness of compression itself. If context length increases infinitely and monotonically, even with kv cache keeping latency and cost low, it is still a growing process and not worth it. What matters more is letting the model know ‘my context is already very long now’, and whether it can offload some information to the file system, just as a person organises memories into documents and puts them in Notion, and next time knows when to go get them back. The model must understand that this thing did not vanish into thin air, it was compressed.
— Zhang TaoPorting a reasoning model to agent scenarios makes results worse
Porting a reasoning model designed for competitive programming or maths directly into agent scenarios makes results worse: instruction following degrades, and the probability of hallucination and hallucinated tool calls rises. The right approach is interleave thinking — following the ReAct framework, after obtaining an observation do not immediately predict the next action, but do a relatively brief stretch of intermediate reasoning, thinking about what has been done and what should be done next. Not like O-series models solving maths problems, where the user gives a very short question and it thinks several thousand tokens all inside its head in a flash.
— Zhang TaoManus is SOTA on Remote Labor Index, with a 2.5% completion rate
scale.ai released a new benchmark called RLI (remote labor index), and Manus is SOTA, first place, beating competitors like Claude and Gemini. The criterion is: could the work this AI system completes make a realistic client willing to pay for it, and is it impossible to tell whether a human or an AI did it. But the completion rate is only 2.5%, far from one hundred percent. He values this benchmark because it fits the metric for a general agent — how much of what a remote worker can do can it complete. Optimistically, maybe by 26 it can be pushed to 20 or 30.
— Zhang TaoDon't hand the model the limits of being human
Many agent companies have inertial thinking, wanting a multi-agent system divided into roles like designer, programmer, manager. He thinks this is wrong: human society has division of labour because no person is all that versatile, and under an organisational structure a great deal of information is lost in communication, adding much friction to cooperation. But a model is something more versatile than a person, so you should fully exploit the model's advantages and not mechanically copy the constraints that come with humans. Building a general agent is building a system that can do what a person does, but you should not demand of the model the division of labour or specialisation of humans.
— Zhang TaoInvite codes were not marketing, cloud vendors said opening up would crash them
After talking to every cloud vendor and inference provider before launch, he was surprised to find that the compute in the world that could be in place immediately the next day was far less than imagined. None of the clouds and model vendors they used could supply that volume; Cloud said you absolutely must not open it up, if you open it up we will crash. So the only option was to control volume, and the way to control volume was invite codes. He admits there were better approaches, such as targeted invitations rather than explicit codes. Less than a month later they removed the invite codes. He says outright: if when we launched in March, if there had been no paid promotion, may my whole family die.
— Zhang TaoMass personalisation need not go through parameter updates
He judges that mass personalization does not necessarily have to be done with online learning or a parametric approach — hanging a set of non-reusable parameters on every user (such as a multi-lora scheme) actually lowers inference efficiency, because cutting cost and latency depends on economies of scale. What is genuinely worth distinguishing is ‘continual learning’ versus ‘online learning’: online learning is only valuable when the ideal distribution of the task changes over time (he cites financial markets, where today's correct answer is not necessarily tomorrow's). Much of what is now called online learning is in essence just on policy data collection plus periodic optimisation, the task itself has no dynamism, and it will quickly be beaten through on benchmarks with no room for continued improvement.
His four suggestions for agentic models
First, rather than infinitely expanding the context window, let the model learn compaction awareness — to realise its context may be compressed, and to know how to offload to and retrieve from the file system. Second, the optimisation target of reasoning should not be made a pure ‘brain in a vat’; it should consider combining observation, i.e. TIR (Tool Integrated Reasoning), which is a different path from solving competitive programming and maths purely with RLVR. Third, interaction mode: in chatbot scenarios the user and the model execute alternately, you wait for me and I wait for you, whereas what was novel when Manus first appeared was that while the agent kept working the user could interject at any time — change the goal, add information, even terminate it — and many models have not yet mastered this asynchronous interaction. Fourth, error resilience: in real environments errors are the norm, and the best model should always be able to find another path to try, which requires dedicated training.
Self-positioning: from the CBA to the NBA
He describes this startup as his first full participation in global competition, joking internally that it is ‘from the CBA to the NBA’. Even though Manus may now already have 100 millionaires (his words), looking horizontally at the top players in every industry he feels he is nothing much, maybe just the NBA average. The cost structure is completely different too: in the mobile internet era you could see IGA behaviour and marginal cost was very low, whereas now from day one of Manus going live it has basically been hundreds of thousands of dollars, hundreds of thousands of dollars burning, a fairly asset-heavy investment from the start.
Technology has a veto over product
Who has more say, product or model? His answer is that technology serves product, but technology has a veto over many things: product may have some very tempting quick-and-dirty approaches, such as whether to abandon the pure sandbox idea and switch to a quick fix, and in such cases both he and the technology side will stand up and stop it outright. He explicitly opposes the concept of voting, believing voting alienates the team — people will serve their own opinions, and you should look at the goal rather than the means of voting; if goals align, consensus can certainly be reached. In discussions no one is domineering, but they expect Red to make the final call; the value of discussion is not to discuss out a result but to provide more alternatives.
Manus's biggest worry is losing its distinctiveness
From the outside, his biggest worry is that Manus loses its distinctiveness; from the inside, what he fears most is that Manus becomes complicated. He quotes the GitHub line — every feature added eats away at everything else — so he hopes Manus can keep walking with restraint, yet restraint must not affect continued growth. As for whether it will die from competition, he thinks the greater possibility for Manus is not losing to competition but users leaving after it loses its unique value, and that also counts as competition, so the answer is: Manus could die because of competition.
Yann LeCun leaving Meta is a positive signal
He thinks Yann LeCun is a respected titan in the industry, but playing such a role inside a commercial organisation has its own pain, and finding free space to do research is good, while it also frees Meta of a lot of ideological baggage, and Meta may invest in some more plain work with quick results. On Tian Yuandong, he particularly champions the latent reasoning direction: RLVR in essence increases the model's stability under pass@1, raising the probability that a single inference reaches the correct answer, but whether the model can solve a problem still depends on base quality — a non-RLVR model sampling multiple times will with high probability also sample a correct trajectory, which shows RLVR is close to solving problems by search, and the sampling step itself brings collapse; latent reasoning does not do this sample, and can consider multiple possibilities simultaneously in near-parallel dimensions, making reasoning more efficient and also solving the problem of high latency from the user's perspective.
In their own words · checked verbatim
But that's what startups are like. One step early and you're a corpse, right? One step early makes you a pioneer, ten steps early makes you a corpse.
但是创业就是这样 你早一步就先裂 对吧 早一步是先驱 早十步就是先裂
Ji Chao23:12
If you finish a product and you don't think it's cool, don't ship it. If you don't think it's cool, nobody will think it's cool.
如果一个产品做完 你觉得不太酷就别发 你都觉得不酷 没人会觉得酷
Ji Chao1:04:30
What's the biggest difference between an agent and a chatbot? In the whole chatbot system there are only two elements: the human user and the model. The two of you interact in a back-and-forth form. But an agent actually has a third element, which is the environment, or the run time.
agent chatbot的最大区别是什么 就是chatbot这整个系统里只有两个元素 人用户以及模型 你们两个之间以往复的形式去交互 但其实agent有一个第三个元素 是环境或者叫run time
Ji Chao1:17:38
It's like part of my memory — like a person, I don't have to keep this thing in my head all the time. Human memory is actually pretty bad, my working memory is very bad, but I'll know that I can organise this thing into a document and put it in my Notion, and next time I'll know when I should go get it back.
就像我有一部分 记忆就像人一样 我这个东西 不用脑子里一直装着 人的记忆其实挺差的 就是我的工作内存很差 但是我会知道 这个事我可以整理成一个文档 放在我的notion里头 我下次我知道 我该什么时候去拿回来
Zhang Tao1:36:46
Actually the model is something more versatile than a person, so you should fully exploit the model's advantages and not mechanically copy the constraints that come with humans.
实际上模型是比人更加全能的一个东西 所以你应该充分利用模型的优势 而不要生搬硬套人带来这套约束
Zhang Tao1:52:53
If when we launched in March, if there had been no paid promotion, may my whole family die.
如果我们在三月份发布的时候 如果没有任何付费宣传我死全家
Zhang Tao2:17:05
The value of discussion, I think, is not to discuss out a result, but for more people to provide more options.
讨论的价值我觉得不是说讨论出一个结果 而是说更多人提供出更多的方案
The greater possibility for Manus is not losing because of competition, but because...
Manus更大的可能性 不是因为竞争而输掉 而是因为
Figures
| Manus's current ARR | Over $100 million | 1:13:36 |
| Manus's input/output token ratio | 100:1 to 1000:1 | 1:31:42 |
| chatbot's input/output token ratio | 3:1 | 1:31:42 |
| Manus's completion rate on the Remote Labor Index | 2.5% | 1:45:51 |
| Manus 1.5's average perceived speedup | 3 to 5 times | 2:29:10 |
| Daily spend from Manus's first day live | Hundreds of thousands of dollars | 3:00:29 |
| Number of Manus millionaires | 100 | 3:00:29 |
| Company size | 100 people | 3:04:33 |
| Largest team previously led | About 10 people | 3:04:33 |
Glossary
- compaction awareness
- The model realising its context may be compressed, and knowing how to offload to and retrieve from the file system.
- Remote Labor Index
- A benchmark released by scale.ai measuring the share of remote-worker tasks an AI completes that a client would be willing to pay for.
- TIR
- Tool Integrated Reasoning — reasoning combined with observation, rather than solving maths problems purely with RLVR.
- RLVR
- Reinforcement learning with verifiable rewards; in essence it increases the model's stability under pass@1.
- latent reasoning
- Not sampling, but considering multiple possibilities simultaneously in near-parallel dimensions, making reasoning more efficient.
How to listen
Founders, investors and engineers building AI applications or agents, especially those concerned with general-agent technical routes, cost structure and product trade-offs.
The chat after 3:18:45 about Yann LeCun, Tian Yuandong and latent reasoning — lower information density.