OpenClaw made Luo Fuli change her verdict: an agent framework can cover a model's weak spots
She first thought it was just an operations-flavored shell. After installing it late one night during Spring Festival, it took over her life and her research within two days. Her conclusion: a complex agent orchestration can lift a mid-tier model's weak spots to near the top.
The video won't play here. Listen to the audio instead:
The argument · tap a timestamp to hear it
It made a skeptic change her verdict within three days
Luo Fuli initially thought OpenClaw was just ‘Claude Code plus an IM UI’, and combined with the founder's operational moves, she was strongly averse to it. She installed it late one night during Spring Festival and talked to it from 2 a.m. until dawn at 6: on day one she felt it had warmth and a soul, and it reminded her to go to bed early; on day two she handed it problems from her life and work, and it solved all of them; on day three she handed it a research task — building a user agent that simulates multi-turn user interaction — and it was done in an hour or two. That is how she saw its product logic clearly: use a whole agent orchestration to cover the model's weak spots, and once you plug it in you no longer have to worry about whether the model has video understanding. In those few days she burned nearly a thousand dollars on day one alone.
— Luo FuliA crowd editing the framework only made it stronger
She pulled her team into a group chat to edit OpenClaw's framework together. In a group of over a hundred people with wildly different backgrounds, all editing frantically, nobody broke the model and nobody broke the agent framework — it actually became smarter. She says if she were editing alone, the framework would progress very slowly; with a crowd editing, it might iterate a round in just a few hours. She connects this with OpenClaw's soaring star count and believes it is something that must happen as a precursor to AGI. The accompanying move was concrete too: she told the team that the next day, anyone with fewer than 100 OpenClaw conversations could just quit — but she doesn't actually evaluate on it, she just wanted to say ‘if you don't use it, you may fall behind’.
— Luo FuliMLA is so perfect there is no room left to play
Kimi and GLM, trained around the same time, both chose MLA. Luo Fuli says MLA is excellent for the Chat era and good for long text too, because it greatly reduces KV Cache. But it doesn't suit the Agent paradigm: MLA was designed from the start to push the ratio of memory access to compute to a perfect critical point, leaving no room to play — if you want to speed it up with speculative decoding like MTP, it gets stuck on compute bound again, so MLA-family models probably all skipped MTP and are somewhat slower. They therefore chose the Hybrid-Tension architecture; the Pro generation pushed the ratio of full attention layers to sliding window layers to 7:1, using sliding windows to save KV Cache and MTP to fill the saved compute back in.
— Luo FuliResearch needs more cards than training does
Her compute allocation is roughly 3:1:1 for research, pre-training and post-training — pre-training and post-training get comparable compute, while research should at minimum exceed the total cards used for formal training; you have to set aside even more cards for research. The real bottleneck is exactly here: ideas are born and code is written too fast, and you get stuck on cards, needing to launch many experiments in parallel to validate one idea. She says a 1T model is the ticket to reaching near-frontier level, training one takes several thousand cards, and the cards actually spent on research are 3 to 5 times that. In the old Chat era the ratio of pre-training to post-training was 3:1 or even 5:1; this year top teams should all be at 1:1.
— Luo FuliHaving a hierarchy means assuming the boss is smarter
Luo Fuli's team is 100 people, with no groups and no job levels; only twenty to thirty or thirty to forty people actually work on iterating one generation of a model, and the intern ratio is high. She explains why no groups: first, many people are interested in both pre-training and post-training, and too-clear group divisions kill creativity; second, people doing pre-training naturally care more about data diversity, and moving to post-training is a good complement. She believes the value of flatness is letting everyone contribute creativity and intelligence equally; any hierarchy is to some degree a set of norms and constraints, and norms and constraints themselves suppress creativity; hierarchy also assumes by default that the person at that level should be smarter than everyone else, which is a very strange definition.
— Luo FuliLarge models have no survival crisis, so they are freer
She puts human evolution and large model evolution into two environments: humans evolved for survival as nature changed, and language came only at the very top, so it is an upright triangle; large models amplified language enormously from the start, so it is an inverted triangle. The difference is that large models have no survival crisis — if they don't replace anyone, they won't die. She judges that without a survival crisis they will instead evolve more freely, more loosely, more creatively, because they have so much compute available, all of human knowledge as a starting point, and so many people helping them improve. This is also why she thinks AI's evolutionary path will not be the same as humans'.
— Luo FuliChina and the US have almost no generation gap in pre-training
She believes several domestic companies already have a 1T base, including KIMI, MIMO and others. At the current level, if reaction speed is fast enough, the generation gap with top models is only two to three months — not catching up to the model of two or three months later, but catching up to the contemporary model. The premise of this judgment is that at the pre-training level there is almost no generation gap, and domestically there may even be an architectural advantage. She also comments on MiniMax: achieving current agent capabilities with a model ten times smaller is impressive, and its post-training agility is very high, but without a 1T base it has not truly matched that level; domestically there is not yet a company with both a base and agility.
— Luo FuliEnvironment buys you capability more than experience does
In her team only about one third to one quarter of people have a bit of training experience, and most have only trained 7B or 14B scale models; large-model experience basically cannot be reused. Her judgment is that these capabilities can be washed away quickly, at most one or two months, or three to four months at the slow end; the key is whether you put people into an environment and drive them with a higher standard. So she says environment matters more than experience; what she cares about is not hiring experienced people but whether she has built an environment where everyone improves faster and learns from each other. This is also why she led a group of people with no large-model background to build Flash: to let this group complete their own evolution in the process of doing the work.
— Luo FuliIn their own words · checked verbatim
It's essentially a middle layer, like a middle between people and the model, right, and this middle layer can be very thick, and instead the front-end UI display is the thinnest layer, it's no longer very critical.
它相当于是一个中间层 它像人和模型之间的中间 对 然后这个中间层 它可以做的非常的厚重 然后反而那个前端的UI展示 它是最薄的一层 它已经不是很关键
Luo Fuli21:09
I think this is also the first time I felt how you use a crowd's wisdom to elevate a thing itself.
我觉得这也是我第一次感受到 你怎么用一群的智慧去提升一个事情 本身
Luo Fuli30:11
Why hasn't it become mainstream? Everyone believes too much in MA. I think everyone believes too much in MA.
为什么它还没有成为一个主流 大家太相信MA了 我觉得 大家太相信MA了
Luo Fuli1:33:09
Because the birth of an idea and, um, getting your hands moving and writing the code out, is too fast, and now what are you stuck on? Stuck on cards.
因为Idea的诞生和这个 嗯 动手 你把它代码写出来 太快了 然后你现在卡在什么呢 卡在卡上
Luo Fuli1:47:07
Any hierarchy should to some degree be norms and constraints, and norms and constraints themselves, I think, suppress creativity.
任何层级应该一定程度上都是在规范和约束 然后规范和约束本身 我自己认为是压制创造力的
Luo Fuli2:00:15
When there is no survival crisis, it instead evolves more freely, and more loosely, more creatively.
当没有生存的危机的时候 它反而会进化的更自由 然后 更散漫 更有创造力
Luo Fuli2:23:29
I think if reaction speed is fast enough, there should be only a two-to-three-month generation gap.
我认为如果反应速度足够快的话 应该只有两三个月的代差
Luo Fuli2:28:30
I think at most one or two months, or three to four months at the slow end, it really can all be washed away quickly, so environment matters more than today.
我觉得最多一两个月 慢的话三四个月 确实都可以被快速洗的 所以环境反而比今天更重要
Luo Fuli3:24:04
Figures
| Single-day Opus 4.6 API spend | nearly $1000 | 23:09 |
| Model inference speed | Flash 100–150 TPS, Pro 60–100 TPS | 1:30:06 |
| Ratio of full attention layers to sliding window layers | 7:1 | 1:30:06 |
| Research : pre-training : post-training compute allocation | about 3:1:1 | 1:48:09 |
| Multiple of research cards relative to training cards | 3–5x | 1:46:09 |
| Training card scale for a 1T model | several thousand cards | 1:45:09 |
| Team size and number actually working on one generation of model iteration | 100 people / twenty to thirty to thirty to forty people | 1:57:14 |
| AGI journey progress and this year's expectation | 20% done, can reach sixty or seventy this year | 2:26:30 |
Glossary
- OpenClaw / open-source Agent framework
- What Luo Fuli calls an epoch-making agent framework, open-source and modifiable, using orchestration to cover the model's weak spots
- MTP / multi-token prediction
- Added during pre-training to improve base capability, and used at inference to fill idle compute and accelerate generation
- MLA / multi-head latent attention
- Written as MA/MIA in the transcript; relies on compressing KV Cache to serve long text, and is said to have no room left for optimization
- Hybrid-Tension / hybrid attention architecture
- Full attention layers and sliding window layers at a 7:1 ratio, saving KV Cache while preserving long-text capability
- Non-Colice / long context (transcript transcription)
- Refers to the efficiency and cost of long context, the core design goal of this generation of model architectures
- Skills
- Task experience accumulated jointly by people and agents, readable by agents, turning tacit experience into reusable assets
How to listen
AI engineers and model team leads, plus investors trying to read the 2026 Agent race and where compute is flowing.
2:19–2:22, the TTS details are mostly product teasers and can be skipped.