Open weights now trail closed models by 3-5 months, and that lag is the whole safety margin
Kimi K3 compresses the open-versus-closed gap from the 6-9 months people had been arguing over down to 3-5 months, and that lag is the only risk buffer anyone has — heavy-handed regulation of open weights would merely postpone the inevitable while voiding the buffer itself.
The video won't play here. Listen to the audio instead:
The argument · timestamps estimated from transcript position
The frontier gap has been compressed to three to five months
K3 is a 2.8T-parameter MoE, released July 16 with the weights promised for July 27. The author pins the entire piece's key conclusion on a single number: the performance gap between open and closed models — or between the US and China — has narrowed from the 6-9 months people had been arguing over to 3-5 months. He adds a robustness note: even if Moonshot never delivers the open weights, as long as China has a closed model of equivalent strength, most of the inferences only land somewhere in the middle, and the direction is unchanged. The leaderboard evidence: second overall on Vals AI, third on the Artificial Analysis intelligence index (losing only to Claude Fable and GPT-5.6 Sol Max, and cheaper than both), first on Frontend Code Arena. This is the closest an open model has come to the frontier since DeepSeek R1.
— Nathan LambertDistillation explains almost none of what K3 achieved
The author states the conclusion outright: if adversarial distillation contributed to K3 at all, it did so only to a very small degree. His reasoning is that R1 and K3 are different in kind — R1 was a Chinese lab pivoting to reasoning models extremely fast and getting there first, while K3 is scaling along already-known directions in data, algorithms, architecture, tooling and environments, which is to say the same set of problems people at OpenAI and Anthropic are working on right now. A second line of evidence comes from his trip to China: Kimi's team culture was one of the few he could feel immediately among the companies he visited, with strong individual execution and freedom of expression despite GPU constraints. Having met them, the result did not surprise him. He expects the distillation argument and the pressure around it to continue, but says observers who attribute Chinese models to IP theft will be proven wrong by reality.
— Nathan LambertThe WAIC speech is Beijing declaring its tolerance for open-weight risk
Until now there had been almost no national-level Chinese policy statement on open-source AI — as far as the author knows, no senior official had commented publicly. That changed this week: in his keynote at the World Artificial Intelligence Conference (WAIC), Xi Jinping (习近平) very directly staked the future of China's AI ecosystem on open source and global diffusion. The author reads it this way: that commitment landing in the same week as the strongest open-weight model in history amounts to an implicit statement of China's risk tolerance for open weights. The simplest explanation is that the Chinese government is tracking model risk closely, possibly at a finer technical granularity than America's "vibes-based regulation," and would act if it measured real risk — and it currently judges that frontier models pose no substantive risk. At the same time, economic policymakers think getting AI adopted first is good: build distribution, monetize later, the same path as cars, solar and advanced manufacturing.
— Nathan LambertOpen weights are decelerationist because they cut frontier labs' margins
Dean Ball — personally a supporter of open source, though now at OpenAI — says open-weight models are fundamentally decelerationist, and the author says he is right. The transmission chain is concrete: strong open models sharply compress the gross-margin room of closed labs, which means both less profit to reinvest in the next generation of models and a market markdown of these companies' terminal values, which in turn discounts future funding rounds. Together, those push back the arrival of the most transformative models. But the author does not think this is enough to knock OpenAI and Anthropic out of the ranks of the world's most valuable companies, and he thinks it is net good for society: on the diffusion side open weights are accelerationist, driving down the entry price of intelligence at a given performance tier and encouraging customization — the curve just starts much more slowly. His phrasing: a slower-starting but potentially larger exponent. The risk is that if closed capability runs too far ahead, customization stops being meaningful.
— Nathan LambertThree architectural changes, not scale alone, bought the 2.5x efficiency
Kimi's launch blog gives verifiable technical handles: K3 is built on two architectural updates, Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), which improve the flow of information along sequence length and along model depth respectively. It also pushes MoE sparsity much higher — 16 of 896 experts actually active — paired with the Stable LatentMoE framework. Together with refinements to the training and data recipe, this delivers roughly a 2.5x improvement in overall scaling efficiency relative to K2. The author also traces how this innovation spread: KDA comes out of the Kimi Linear paper and shares lineage with the Gated DeltaNet he used while building Olmo Hybrid at Ai2; Gated Delta Networks were only proposed in late 2024, building on ideas from Mamba, and by mid-2026 they are inside frontier models, with Qwen's latest model having switched to a related architecture too.
— Nathan LambertChinese labs are more capital-efficient because catching up is cheaper
The author judges that Chinese labs are clearly more capital-efficient, and in a world where intelligence scales with effective capital — buying compute, data and talent — that may be the single largest advantage an AI industry can have. He lists several factors he has heard but cannot confirm: Chinese researchers are paid less yet are better at solving LLM research puzzles; a data industry is taking shape (with budgets still far below Anthropic's billion-dollar scale); meaningful training compute is being obtained by working around export controls; and inference demand is far smaller, so more compute can go into training — though Moonshot has already paused new subscriptions because of K3, while the API keeps running. His mechanistic explanation: American labs spend enormous effort on dramatic leaps that push the frontier, whereas Chinese labs are more focused on catching up, and catching up is itself cheaper — in the same way a student model in distillation can surpass its teacher.
— Nathan LambertChina is not maintaining its open ecosystem, it is doubling down
The weekend after K3 shipped, Alibaba announced that Qwen 3.8, at 2.4 trillion parameters, would be released with open weights. The author stresses what that signals: historically Alibaba kept its largest models exclusive to its cloud business as APIs, so the reversal shows Chinese companies are not merely maintaining the open-source status quo but adding to it — and the directional point holds even if Qwen 3.8 scores below K3. He offers a checkable prediction: if Qwen 3.8 ships before the next generation of Gemini, it pushes Google down to eighth place on the list of labs with the smartest models. Separately, DeepSeek V4 is expected to graduate from preview to a full release. In his ranking of each lab's peak model, DeepMind is carried only by Gemini Flash 3.5; he describes seeing DeepMind and several American giants sitting that low as shocking, and argues the X AI team deserves more credit than it gets.
— Nathan LambertRestricting open models would leave American defenses asymmetrically exposed
The author's position: if Claude Mythos were open-sourced today, the negative consequences would be relatively mild and the risks are overhyped — but that will not hold forever, so having the strongest models available in controlled, closed form a few months ahead of open models at the same level is a very safe equilibrium. The danger is regulation going wrong. He cites Axios's reporting: the Commerce Department last year considered putting several Chinese AI labs on the Entity List; the NSA and the White House Office of the National Cyber Director considered issuing a threat advisory aimed at Chinese labs; the White House considered an executive order requiring that US companies host Chinese models only if they can guarantee safety and accept liability for leaks; and Commerce also drafted rules targeting Chinese open models under its supply-chain security authorities. The result would be a highly asymmetric state: America's best models operating with guardrails on cybersecurity tasks, while actors worldwide use excellent Chinese open models to probe American defenses.
— Nathan LambertIn their own words · checked verbatim
The key fact is that either the open-to-closed or American-to-Chinese model performance gap has been reduced from the debated 6-9 months to something shorter, say 3-5 months.
Nathan Lambert0:00
Open-weight models are inherently decelerationist, and I’m continually surprised to see the so-called “accelerationists” so excited about open-weight models.
Dean Ball0:00
It is becoming clear that the Chinese labs are far more capital efficient. In a world where scaling laws dictate that intelligence is proportional to effective capital – which buys compute, data, & talent – that may be the greatest strength your AI industry could ever have.
Nathan Lambert9:23
Just as the student model can outperform the teacher in distillation generally (not limited to the adversarial distillation of the Chinese labs), an approach of “trying to catch up” rather than “invent the next paradigm” could lead to stronger models.
Nathan Lambert9:23
If we regulate open-weight models heavy-handedly, I suspect much of the world will be lulled into thinking we no longer need to act. All we would’ve done is slightly delayed the inevitable — open models will continue to cross all the key capability thresholds eventually and regardless of legality.
Nathan Lambert19:54
Figures
| K3 release date / weights release date | Released July 16, weights promised for July 27 | 0:00 |
| Lag of open weights behind closed models | Narrowed from the 6-9 months under debate to 3-5 months | 0:00 |
| K3 leaderboard placements | 2nd on Vals AI, 3rd on Artificial Analysis, 1st on Frontend Code Arena | 0:00 |
| Compute available per OpenAI researcher (author's joking estimate) | Thousands of H100-equivalent machines | 0:00 |
| K3's MoE activation ratio | 16 of 896 experts active | 9:23 |
| K3's scaling efficiency gain over K2 | About 2.5x | 9:23 |
| Qwen 3.8 parameter count | 2.4 trillion parameters, to be released with open weights | 9:23 |
Glossary
- Kimi Delta Attention (KDA)
- An attention variant from the Kimi Linear paper that improves how information flows along sequence length.
- Attention Residuals (AttnRes)
- An architectural change that improves how information flows through the depth of the model.
- Stable LatentMoE
- The mixture-of-experts framework used alongside extremely high sparsity (16 of 896 experts).
- adversarial distillation
- Extracting capability back out of a closed frontier model's outputs in order to train your own model.
- decelerationist
- A force that slows the pace of AI investment by depressing frontier labs' profits and their fundraising.
- Entity List
- A US Commerce Department control list; parties on it need licenses to obtain American technology and services.
How to listen
Investors tracking the US-China model landscape, engineering leads who have to choose between open weights and closed APIs, and anyone doing analysis on AI export controls and open-source policy.
Section 5 on US regulatory moves — skippable if policy is not your concern.