The world is too loud. Read what matters.

张小珺·商业访谈录

Pure linear attention doesn't work; a 3-to-1 mix is the emerging consensus

Pure Linear Attention has a fundamental flaw: no guaranteed floor. Kimi and Qwen have both converged on inserting 1 global attention layer for every 3 linear layers, while MiniMax retreated from M1 to Full Attention because its evaluation pipeline missed multi-hop reasoning.

Linear AttentionModel ArchitectureKimiDeepSeekInference Efficiency
Hardcore architecture talk. It explains Kimi Linear's KDA module, where the 3-to-1 ratio comes from, why MiniMax went back to Full Attention, and how linear and sparse attention can be combined.

The argument · tap a timestamp to hear it

12:14

Long chains of thought are blowing up decoding cost

The direct trigger for Kimi redoing its attention mechanism early this year was the long chains of thought brought by the RL wave around DeepSeek R1 and Kimi 1.5, which often run to tens of thousands of tokens. If every layer is quadratic attention, decoding has to store a large KV cache per layer, and the time complexity of decoding L tokens is also quadratic — far too expensive. Hybrid attention can push inference cost down a lot, which is very useful in the context of long chains of thought and agentic AI.

— Yang Songlin
15:14

KDA swaps coarse-grained decay for fine-grained decay

The linear module Kimi Linear chose is called KDA (Kimi Delta Attention), an improvement on Yang Songlin's Gated DeltaNet from last year. Gated DeltaNet was limited by efficiency at the time and used coarse-grained gating similar to Mamba2 — all dimensions under one attention head shared the same decay rate. KDA replaces that with an independent decay rate per dimension, so the RNN hidden state corresponding to each dimension has its own update-like frequency. Intuitively this makes better use of the limited hidden state and improves performance.

— Yang Songlin
18:25

Kimi clears levels internally with a Scaling Ladder

Kimi has an internal mechanism called Scaling Ladder: if it performs well at one scale, it has to keep scaling at the next scale, clearing levels one by one against Full Attention like beating a game. Zhang Yu first tried a whole bunch of ways to mix Gated DeltaNet and found that mixing Gated DeltaNet was better than the alternatives, but in some places it still fell short of Full Softmax Attention; later, swapping in finer-grained decay brought a fairly large improvement.

— Yang Songlin
23:28

Why MiniMax retreated to Full Attention

MiniMax M1 was a very large-scale hybrid attention deployment. On the monitored metrics Lightning Attention performed well and was more efficient, so they shipped it. But they later found that on multi-hop reasoning tasks the drop was very large — the original evaluation pipeline only looked at capabilities like MMLU and did not test multi-hop reasoning. Yang Songlin thinks Lightning Attention is a relatively weak linear attention whose mechanism looks like it stalled two years ago. M2 retreated to full Softmax Full Attention, but they say they are still exploring hybrid attention, and the next version, M3, might switch back again.

— Yang Songlin
38:55

Pure linear doesn't work; only a hybrid has a floor

The current consensus is that pure Linear Attention doesn't work, and has a fairly fundamental flaw on long text: the number of RNN states is constant, so as context length grows it will eventually fail to hold everything and lose precision. Hybrid attention keeps a certain number of global attention layers, so the floor is guaranteed, and handling long text is guaranteed too. Kimi Linear and Qwen 3 Next showed no drop on tasks like ruler. But Yang Songlin also says he doesn't know whether this counts as consensus, because DeepSeek is still trying sparse attention approaches.

— Yang Songlin
40:56

3-to-1 is fast becoming the consensus mixing ratio

Kimi Linear inserts 1 full attention layer for every 3 KDA layers. Yang Songlin says 3-to-1 is now almost becoming a formula: MiniMax was previously 7-to-1, with too few softmax layers, so the long-text guarantee wasn't as good; ByteDance published a paper studying how much softmax ratio a hybrid architecture needs, ran many pre-train-from-scratch experiments, and also concluded that 3-to-1 is best and the Gated DeltaNet module is better than other candidates. Qwen 3 Next also uses 3-to-1 with Gated DeltaNet.

— Yang Songlin
49:01

The ideal architecture replaces global layers with sparse ones

Yang Songlin wrote on Zhihu: why not combine the two approaches? Let sparse attention replace the global attention layers in hybrid attention, so you no longer need the complexity of global attention, but you still have to store KV cache, and the KV cache of the many remaining layers can be shrunk via linear attention. He thinks linear attention's competitor is more Sliding Window Attention than sparse. As far as he knows, no one in industry has combined sparse and linear at the same time; there is some exploration in academia.

— Yang Songlin
1:07:24

China's algorithmic innovation is stronger because it has fewer GPUs

Yang Songlin judges that China's algorithmic innovation is definitely stronger, and on infra architecture China is stronger too. The reason is a different position in the ecosystem: China doesn't have that many GPUs, so the demand for efficiency is higher and there is more incentive to try efficient Linear Attention variants; some Silicon Valley companies have too many GPUs and buy them at high prices. He also says US companies put more into optimizers (optimization), and Kimi was one of the first places to eat the Muon crab.

— Yang Songlin
1:39:37

Hardware is racing toward ever-faster matrix multiplication

Yang Songlin says hardware and the Transformer are co-evolving, and hardware is turning into a shape the Transformer likes better: Tensor Core, TMA, and the separate memory for matrix multiplication on Blackwell are all there to optimize matrix multiplication. Even FA4, because matrix multiplication got so fast, made the exp module the bottleneck, and they use an approximate method to compute exp. So algorithm design must satisfy general hardware principles, otherwise in today's scalability scenarios it has basically no practical value.

— Yang Songlin

In their own words · checked verbatim

I think the consensus now is that pure linear attention doesn't work.

我觉得现在共识的时候就是说 纯linear attention是不work的

Yang Songlin38:55

So you definitely have to — first you have to make your algorithm satisfy some very general principles.

那你肯定你要 首先你要 啊 让你的算法 先去 满足一些 非常通用的原则嘛

Yang Songlin1:34:37

Figures

Kimi Linear mixing ratio1 full attention layer inserted for every 3 KDA layers40:56
MiniMax M1 mixing ratio7 to 140:56
Qwen 3 Next RoPE ratio25% RoPE, 75% NoPE1:11:27
Kimi's RoPE ratiocut by 100%1:11:27
Long chain-of-thought lengthtens of thousands of tokens12:14
Training length in the BERT era51234:51
Long text as seen at the time819234:51
Example forget rate for input-independent decay0.9928:39
Earliest year DeltaNet appeared20211:17:28
Fine-grained decay can be traced back to20161:25:33

Glossary

Linear Attention
An attention mechanism that drops softmax and cuts complexity from quadratic to linear, and can be written in an RNN inference form.
KDA / Kimi Delta Attention
The linear attention module used by Kimi Linear, based on Gated DeltaNet, replacing coarse-grained decay with per-dimension independent decay.
Delta Rule
Use the key to retrieve the old value, linearly combine it with the input value, and write it back to the memory network, including a subtractive erase operation.
Sparse Attention
Only the Top-K tokens participate in attention computation, reducing the cost of generating each token, but it does not save KV cache.
Sliding Window Attention
Attention that only looks at a local window, with the KV cache upper bound bounded by the window size.
Scaling Ladder
An internal Kimi mechanism: if it performs well at one scale, it goes to the next scale and keeps scaling, like clearing levels.

How to listen

Who it's for

Engineers, researchers and investors watching LLM architecture and inference efficiency, especially anyone trying to sort out the Linear Attention vs. Sparse Attention route debate.

Skip

After 1:15, the chat about personal PhD experience and archaeological method can be fast-forwarded.