Chinese open-weight models lead, and distillation explains only one to two months of the gap
Chinese models lead across the board on downloads and benchmarks, but distillation explains only one to two months of the gap; the real reasons are release cadence and task focus, and US companies are already building products on Chinese models.
The video won't play here. Listen to the audio instead:
The argument · timestamps estimated from transcript position
Open weights are not the same as open source
An open-weight model releases only the weights and a license; an open-source model also releases the training code and data. America's genuinely open-source models are led by nonprofits, such as the Allen Institute for AI's Olmo, OpenAthena's Marin, and EleutherAI's Pythia. Nvidia's Nemotron is more open than most open-weight models, having released a large amount of training data, but because it has not released all of it, it still does not count as fully open source. The distinction matters: when people discuss "open source," what they are actually talking about is mostly open weights.
Chinese models have crossed the agentic threshold
GLM-5.2 and Kimi K3 achieved a step change in commercial viability, crossing an agentic capability threshold similar to the one Anthropic's Claude Code crossed in December 2025. Chinese open-weight models passed the US in these two key areas roughly 18 months ago. On Hugging Face downloads, China took the lead in July 2025, mainly on the back of Alibaba's Qwen. Data from the ATOM project, which the author maintains, shows China's download lead has grown to about 1.6 billion out of a total of 3.2 billion, twice that of the US.
The benchmark score gap is stark
On the Artificial Analysis Intelligence Index, as of September 14, 2026, the top three are Chinese models: Z.ai's GLM-5.3 and GLM-5.3-Flash score 45 and 42, and Moonshot AI's Kimi K3 scores 44. The US leaders are Thinking Machines' Inkling and Inkling Small, both scoring 26, and Nvidia's Nemotron 3 Ultra at 23. The top US models were released in June and July 2026, and update less frequently than their Chinese counterparts. Chinese labs are closest on tasks with well-defined demand, such as agentic coding, and lag further behind on open-ended scientific tasks like physics and biology.
Distillation explains only one to two months of the gap
Distillation is a technique Chinese labs use often, training their own models on the outputs of stronger models, but the author estimates that even if distillation were blocked entirely, for instance by deploying KYC tools at Anthropic and OpenAI, the gap from the strongest US models to Chinese open-weight models would widen by only 1-2 months. Distillation has the biggest effect in new domains, and it cannot easily produce a comprehensively strong final model. Chinese labs release faster and have a narrower task distribution, which slightly inflates scores on public benchmarks, because all labs keep improving and later-released models naturally score higher.
US companies are building products on Chinese models
Harvey, Cursor and DoorDash use Kimi, Airbnb uses Qwen, and Perplexity uses DeepSeek. A large number of young Silicon Valley startups use Chinese models to get a low-cost, flexible option. US startups sign enterprise agreements with Chinese model labs to license products for use, a new form of cross-border technology cooperation the author has not seen in his career. On OpenRouter, the Chinese model share has grown from about 70% to more than 80%, and weekly tokens processed have grown from about 1 trillion in September 2025 to about 80 trillion today. On OpenCode, Chinese models account for about 95% or more of inference volume.
Academia has already been taken over by Qwen
The author scanned papers in the five most popular ML categories on arXiv (cs.AI, cs.CL, cs.CV, cs.LG, stat.ML). In January 2023, papers mentioning any open model were 2%; by September 2026 they reached 50%. In April-May 2023, of about 12,000 new papers, 2,600 mentioned at least one open-model family, 5.5% mentioned Llama, and 1% mentioned Chinese models. At Llama's peak in fall 2024, 23% mentioned Llama and 7.5% mentioned Qwen. Today Llama is still mentioned by about 21% of papers, while Qwen has risen to 30%. Any Chinese open-weight model is mentioned by more than 40% of papers, versus 30% for the US.
US models see higher adoption at equal capability
Research shows that US models see disproportionately high adoption when capability and scale are comparable. OpenAI's gpt-oss, its first open-weight models since ChatGPT, became one of the most widely adopted open-weight models ever. Google's Gemma 4 is one of the few models whose adoption approaches Qwen's most popular small models. Nvidia's Nemotron, despite being more capable at the same scale, has seen middling adoption. This shows US models have brand and ecosystem advantages, but not enough to reverse the overall trend.
Open-model risk requires ecosystem-level preparation
The structural challenge of open-weight models is that there is almost no effective way to stop pieces of open-source software from falling into the wrong hands. If access to China's strongest open-weight models were restricted because of risk, the ones hurt would be US companies. HuggingFace once used a Chinese open-weight model to understand a cyberattack, because closed models refused the request. Managing open-weight model risk often comes down to ecosystem preparation. Open models are becoming an essential tool for AI diffusion, and continuing to invest in US open models is the best path to staying ahead of these risks.
In their own words · checked verbatim
Since about April 2025, Chinese AI companies have been the clear leader in open-weight models.
I estimate that if distillation was fully prevented, e.g. with know-your-customer (KYC) tools at Anthropic and OpenAI, the gap from the strongest American models to Chinese open-weight models would only increase by 1-2 months.
There is a growing trend of American startups and companies entering enterprise agreements with Chinese model labs in order to get permission to use their models in their products – a new form of cross-border technology collaboration I have not witnessed in my career.
To a first order approximation, most of academic research is conducted on Alibaba’s Qwen family of models.
In our research , we find that American models of comparable capabilities-to-size regions to their Chinese counterparts get adopted at disproportionate rates.
If an attempt was made to restrict access to the strongest open-weight models from China because they amplify risks, the parties who would be set back are American businesses.
Open models are going to be the substrate for everyone else in the world outside of the few true frontier AI labs, to harness an acceleration in software engineering and other computational practices.
Figures
| GLM-5.3 AAII score | 45 | 0:00 |
| Kimi K3 AAII score | 44 | 0:00 |
| Thinking Machines Inkling AAII score | 26 | 0:00 |
| Nvidia Nemotron 3 Ultra AAII score | 23 | 0:00 |
| Chinese model share on OpenRouter | from about 70% to more than 80% | 0:00 |
| Share of arXiv papers mentioning open models | 2% in January 2023, 50% in September 2026 | 0:00 |
| Share of arXiv papers mentioning Qwen | 30% | 0:00 |
| Share of arXiv papers mentioning any Chinese open-weight model | more than 40% (US 30%) | 0:00 |
Glossary
- open-weight
- Model weights are publicly downloadable, but the training code and data are not necessarily public.
- distillation
- Training one model on the outputs of a stronger model to get comparable capability at lower cost.
- ATOM Project
- The American Truly Open Models project started by the author, tracking data on US open models.
- AAII
- A popular model capability benchmark that aggregates scores across many tasks.
- KYC
- Restricting API access through identity verification, to prevent models from being used for distillation and similar purposes.
How to listen
Founders, investors and policy researchers watching US-China AI competition and the open-model ecosystem, and anyone who wants to grasp the 2026 open-model landscape quickly.