Pure linear attention doesn't work; a 3-to-1 mix is the emerging consensus
Pure Linear Attention has a fundamental flaw: no guaranteed floor. Kimi and Qwen have both converged on inserting 1 global attention layer for every 3 linear layers, while MiniMax retreated from M1 to Full Attention because its evaluation pipeline missed multi-hop reasoning.
The video won't play here. Listen to the audio instead:
The argument · tap a timestamp to hear it
Long chains of thought are blowing up decoding cost
The direct trigger for Kimi redoing its attention mechanism early this year was the long chains of thought brought by the RL wave around DeepSeek R1 and Kimi 1.5, which often run to tens of thousands of tokens. If every layer is quadratic attention, decoding has to store a large KV cache per layer, and the time complexity of decoding L tokens is also quadratic — far too expensive. Hybrid attention can push inference cost down a lot, which is very useful in the context of long chains of thought and agentic AI.
— Yang SonglinKDA swaps coarse-grained decay for fine-grained decay
The linear module Kimi Linear chose is called KDA (Kimi Delta Attention), an improvement on Yang Songlin's Gated DeltaNet from last year. Gated DeltaNet was limited by efficiency at the time and used coarse-grained gating similar to Mamba2 — all dimensions under one attention head shared the same decay rate. KDA replaces that with an independent decay rate per dimension, so the RNN hidden state corresponding to each dimension has its own update-like frequency. Intuitively this makes better use of the limited hidden state and improves performance.
— Yang SonglinKimi clears levels internally with a Scaling Ladder
Kimi has an internal mechanism called Scaling Ladder: if it performs well at one scale, it has to keep scaling at the next scale, clearing levels one by one against Full Attention like beating a game. Zhang Yu first tried a whole bunch of ways to mix Gated DeltaNet and found that mixing Gated DeltaNet was better than the alternatives, but in some places it still fell short of Full Softmax Attention; later, swapping in finer-grained decay brought a fairly large improvement.
— Yang SonglinWhy MiniMax retreated to Full Attention
MiniMax M1 was a very large-scale hybrid attention deployment. On the monitored metrics Lightning Attention performed well and was more efficient, so they shipped it. But they later found that on multi-hop reasoning tasks the drop was very large — the original evaluation pipeline only looked at capabilities like MMLU and did not test multi-hop reasoning. Yang Songlin thinks Lightning Attention is a relatively weak linear attention whose mechanism looks like it stalled two years ago. M2 retreated to full Softmax Full Attention, but they say they are still exploring hybrid attention, and the next version, M3, might switch back again.
— Yang SonglinPure linear doesn't work; only a hybrid has a floor
The current consensus is that pure Linear Attention doesn't work, and has a fairly fundamental flaw on long text: the number of RNN states is constant, so as context length grows it will eventually fail to hold everything and lose precision. Hybrid attention keeps a certain number of global attention layers, so the floor is guaranteed, and handling long text is guaranteed too. Kimi Linear and Qwen 3 Next showed no drop on tasks like ruler. But Yang Songlin also says he doesn't know whether this counts as consensus, because DeepSeek is still trying sparse attention approaches.
— Yang Songlin3-to-1 is fast becoming the consensus mixing ratio
Kimi Linear inserts 1 full attention layer for every 3 KDA layers. Yang Songlin says 3-to-1 is now almost becoming a formula: MiniMax was previously 7-to-1, with too few softmax layers, so the long-text guarantee wasn't as good; ByteDance published a paper studying how much softmax ratio a hybrid architecture needs, ran many pre-train-from-scratch experiments, and also concluded that 3-to-1 is best and the Gated DeltaNet module is better than other candidates. Qwen 3 Next also uses 3-to-1 with Gated DeltaNet.
— Yang SonglinThe ideal architecture replaces global layers with sparse ones
Yang Songlin wrote on Zhihu: why not combine the two approaches? Let sparse attention replace the global attention layers in hybrid attention, so you no longer need the complexity of global attention, but you still have to store KV cache, and the KV cache of the many remaining layers can be shrunk via linear attention. He thinks linear attention's competitor is more Sliding Window Attention than sparse. As far as he knows, no one in industry has combined sparse and linear at the same time; there is some exploration in academia.
— Yang SonglinChina's algorithmic innovation is stronger because it has fewer GPUs
Yang Songlin judges that China's algorithmic innovation is definitely stronger, and on infra architecture China is stronger too. The reason is a different position in the ecosystem: China doesn't have that many GPUs, so the demand for efficiency is higher and there is more incentive to try efficient Linear Attention variants; some Silicon Valley companies have too many GPUs and buy them at high prices. He also says US companies put more into optimizers (optimization), and Kimi was one of the first places to eat the Muon crab.
— Yang SonglinHardware is racing toward ever-faster matrix multiplication
Yang Songlin says hardware and the Transformer are co-evolving, and hardware is turning into a shape the Transformer likes better: Tensor Core, TMA, and the separate memory for matrix multiplication on Blackwell are all there to optimize matrix multiplication. Even FA4, because matrix multiplication got so fast, made the exp module the bottleneck, and they use an approximate method to compute exp. So algorithm design must satisfy general hardware principles, otherwise in today's scalability scenarios it has basically no practical value.
— Yang SonglinIn their own words · checked verbatim
I think the consensus now is that pure linear attention doesn't work.
我觉得现在共识的时候就是说 纯linear attention是不work的
Yang Songlin38:55
So you definitely have to — first you have to make your algorithm satisfy some very general principles.
那你肯定你要 首先你要 啊 让你的算法 先去 满足一些 非常通用的原则嘛
Yang Songlin1:34:37
Figures
| Kimi Linear mixing ratio | 1 full attention layer inserted for every 3 KDA layers | 40:56 |
| MiniMax M1 mixing ratio | 7 to 1 | 40:56 |
| Qwen 3 Next RoPE ratio | 25% RoPE, 75% NoPE | 1:11:27 |
| Kimi's RoPE ratio | cut by 100% | 1:11:27 |
| Long chain-of-thought length | tens of thousands of tokens | 12:14 |
| Training length in the BERT era | 512 | 34:51 |
| Long text as seen at the time | 8192 | 34:51 |
| Example forget rate for input-independent decay | 0.99 | 28:39 |
| Earliest year DeltaNet appeared | 2021 | 1:17:28 |
| Fine-grained decay can be traced back to | 2016 | 1:25:33 |
Glossary
- Linear Attention
- An attention mechanism that drops softmax and cuts complexity from quadratic to linear, and can be written in an RNN inference form.
- KDA / Kimi Delta Attention
- The linear attention module used by Kimi Linear, based on Gated DeltaNet, replacing coarse-grained decay with per-dimension independent decay.
- Delta Rule
- Use the key to retrieve the old value, linearly combine it with the input value, and write it back to the memory network, including a subtractive erase operation.
- Sparse Attention
- Only the Top-K tokens participate in attention computation, reducing the cost of generating each token, but it does not save KV cache.
- Sliding Window Attention
- Attention that only looks at a local window, with the KV cache upper bound bounded by the window size.
- Scaling Ladder
- An internal Kimi mechanism: if it performs well at one scale, it goes to the next scale and keeps scaling, like clearing levels.
How to listen
Engineers, researchers and investors watching LLM architecture and inference efficiency, especially anyone trying to sort out the Linear Attention vs. Sparse Attention route debate.
After 1:15, the chat about personal PhD experience and archaeological method can be fast-forwarded.