The world is too loud. Read what matters.

张小珺·商业访谈录

DeepSeek trained a 671B model on 2000 GPUs by betting on techniques no one dared try

Every DeepSeek innovation was forced out by cost: from MoE to MLA to FP8 training, they tried what others wouldn't at scale — and finished a single training run with no rollback.

DeepSeekLarge ModelsMoEReinforcement LearningPaper Breakdown
A paper-by-paper breakdown of DeepSeek's nine papers, explaining the technical motivations and engineering details behind MoE, MLA, FP8, MTP and R1's reward design. High information density, suited to anyone who wants to understand DeepSeek's technical路线.

The argument · tap a timestamp to hear it

0:00

Innovation was forced out by cost

He Junxian argues DeepSeek stopped purely following others' practice very early on — Llama and Mistral were both authoritative, but starting with DeepSeek MoE they explored more experts. Not innovation for its own sake, but a real desire to push down cost and make the whole thing more efficient, which then diverged from others and grew increasingly different. This thread runs through the whole piece: from MoE to V2's MLA, every innovation has a cost motive behind it.

— He Junxian
5:05

The reward-model path took a half-year detour

He Junxian's own team was doing similar work more than two months before R1's release, tried many complicated things that didn't work, including using a reward model for reinforcement learning, with results that were never ideal and repeated setbacks. In the end they found the simplest thing worked best: only rule-based reward, checking whether the final answer to a math problem is correct, running unit tests for coding, no need for another model to judge whether the generation is right. He thinks this detour had a lot to do with OpenAI's early paper on process-supervision reward models, which the whole community followed, but turns out what OpenAI itself ended up doing may not have been that.

— He Junxian
8:07

High-Flyer opened 5000 A100s to universities for free

He Junxian was already watching High-Flyer in 2022, before DeepSeek was founded. High-Flyer advertised 5000 Nvidia A100s at the time, a huge number then, because ChatGPT didn't exist yet and no one had the concept of large models. High-Flyer couldn't use that much compute itself, so it built a cluster and a scheduling software system and opened it free to university researchers. He Junxian later tried it and found it very impressive: High-Flyer itself was tiny, maybe just over 100 people, yet used 5000 cards to build a very mature supercomputing system, called in Chinese the Firefly cluster.

— He Junxian
38:37

DeepSeek publicly admitted it could game benchmarks

The C-Eval He Junxian's team built was the first Chinese benchmark for large models in May 2023. Crazy benchmark gaming appeared quickly, followed by high scores with low ability; companies gamed benchmarks more or less but no one talked about it. DeepSeek instead wrote it into a paper with a rigorous controlled experiment: just train on lots of multiple-choice questions and the benchmark can jump 20-plus points instantly, e.g. from 47 to 71. He Junxian says this table, though on page 21 near the back of the paper, was an important result for understanding benchmark gaming at the time. He later did new evaluation and found DeepSeek's released base model hadn't gamed benchmarks, while quite a few domestic models with high scores actually had. Kunlun Wanwei's Skywork also wrote a similar controlled experiment around the same time.

— He Junxian
1:07:02

MLA is the first thing DeepSeek proposed itself

He Junxian distinguishes two kinds of innovation: adding shared experts, increasing the number of experts — still improvements on others' work; but Multi-head Latent Attention was first proposed by DeepSeek, not a modification of prior work. It jointly compresses K and V into a low-dimensional latent vector; at deployment only this compressed latent is stored, not K and V directly, cutting KVCache by 93%, roughly one-tenth of before. The cost is that it's more aggressive than GQA, but the effect is much better than GQA with only 2.25 groups. He Junxian says when it first came out people didn't yet see it as so fundamental; only after DeepSeek became hot did many people abroad start reading these papers and paying attention.

— He Junxian
1:36:18

The V2 price war: not losing money, but making it

He Junxian recalls that after DeepSeek V2 launched in May 2024 it triggered a price war in China, with one yuan or a few yuan per million tokens, far below OpenAI and several times below other domestic vendors, possibly not even the same order of magnitude. The key is he heard DeepSeek wasn't deploying at a loss — it was actually still profitable, just not by much. He was shocked: deploying a very large model on relatively poor GPUs at a very cheap price, and still making a profit. He remembers people starting to call DeepSeek the Pinduoduo of large models.

— He Junxian
1:42:25

V3 finished in one training run, never rolled back

DeepSeek V3 specifically wrote a line in its abstract: the entire training process saw no loss spike, it trained through in one go, without any rollback. This is extremely rare in pretraining — in the past GPU failures and unknown causes would make loss suddenly spike, forcing a stop and rollback to retrain. So V3's paper devotes a lot of space to engineering implementation, a different style from earlier papers; the author sees this as flexing muscle, showing DeepSeek's infra is exceptionally good.

— narration
1:43:28

2000 H800s trained a 671B model

V3 is 671B parameters, 14.8T tokens, using only 2000 H800s, not H100s. The author stresses this scale isn't large for a 600-plus-B model; at the time many companies at home and abroad already had over ten thousand cards, and abroad even hundreds of thousands of H100s or better. The equivalent price was 5.57M USD, over five million dollars, shocking at the time. Compared with Llama 3's 400B model, which the author remembers cost 30 million dollars, that's roughly a 6x gap, and possibly not even 30 million.

— narration
1:46:28

MTP is a technique no one dared use at scale

V3 added multi token prediction: during training it predicts not just the next word but the next three words at once, the benefit being a denser training signal, and the model may learn to plan further ahead for tokens. The idea comes from a not particularly famous paper; no one had really used it at very large scale before, and DeepSeek itself hadn't used it in V2. The author considers this still a very brave act, because once training becomes unstable or an unexpected issue appears, the whole thing changes; an ordinary team might not have the atmosphere to do it.

— narration
1:53:31

FP8 training: almost no one had done it successfully

V3 used fp8 training, i.e. low precision training, where many intermediate vectors aren't represented in 32-bit or 16-bit floating point, making training faster and more efficient. But the challenge is big: decimals are imprecise, training may not work, may be unstable, or the effect may worsen. The author says that although many people do quantization and low precision at deployment, almost no one had successfully done it in real large-scale training; DeepSeek V3 may be the first or among the earliest to really validate and successfully do mixed-precision training on a large-scale language model. They did very careful controlled experiments and found some intermediate variables still had to use original precision for training to work.

— narration
1:57:33

30B active params vs 400B: 10x deployment cost gap

Except for its first large model, DeepSeek has been MoE from DeepSeek MOE to V2 to V3; Llama 3 went from Llama 1 to Llama 3 wearing such a large Dense model, 400B still not MoE, leading to particularly high training cost and particularly many active parameters. DeepSeek V3 has only 30B active parameters, meaning deployment cost is more than 10x less than Llama 3's 400B model. On effect, English is comparable to Llama 3, while on reasoning, Code, Math and Chinese it greatly exceeds Llama 3's 400B Base.

— narration
2:40:56

RL may just rank the correct answer first

Figure 7 of DeepSeek Math gives a negative signal: when k equals 1, the blue line with RL is several points higher than the green line without RL, looking like RL works well; but as k increases and multiple responses are sampled, on simple math levels the green line even ends up higher, showing that after RL the model's exploration ability actually declined. Their own bolded conclusion is the improvement is attributed to boosting the correct response, not enhancement of fundamental ability. The author says no one made this observation at the time; DeepSeek, despite its own results being much better, poured cold water on itself, very rigorous.

— narration
2:42:56

Rule-based reward is more robust than a reward model

DeepSeek Math discusses how to achieve more effective RL and mentions improving the reward model's generalization, but the author thinks from today's view there's another path: don't use a reward model, only rules, and it's the most robust. Because the rule of whether a math final answer is right or wrong is universal — whether high school, elementary, middle school or university problems, if the standard answer is matched the response is basically considered correct; whereas a reward model trained on elementary and middle school data judges those accurately but may suddenly judge university math problems inaccurately. However, open-ended domains, problems without a concept of right or wrong, still need a reward model.

— narration
2:55:00

R1 dropped the reward model, contradicting Coder V2

R1's reward has only two: accuracy reward checking correctness, format reward checking whether the output follows the desired format, both rule-based, no reward model. The author points out this contradicts earlier results: in DeepSeek Coder V2 they still had experiments saying using a Reward Model for Code was better and Compiler worse, but R1 on LiveCode only uses Compiler, not Reward Model. From DeepSeek Prover onward no reward model was used, and by R1 the reward model was gone.

— narration
3:18:23

Process-supervision reward model PPO tuned for two or three months without success

The first time doing long-chain reasoning, the team directly took DeepSeek's open-source process-supervision reward model (similar to the DeepSeek Math set) to do online PPO, worked for a long time, tuned for two or three months, results were never good, and finally gave up reinforcement learning to do iterative self-evolution, which instead worked. After the project finished they still judged it wasn't enough, still needed to do online, so at the end of last year they restarted PPO.

— Guest
3:19:23

Only after dropping the reward model did PPO work easily

The key difference that made the redo of PPO work at the end of last year was: no reward model. The guest says this matches the route everyone converged on — people found using a reward model made things somewhat difficult, then stopped using it, and DeepSeek and R1 both converged here. In between they first did a version of rejection-sampling SFT that was easier to make work, which went relatively smoothly, but the team believed the self-evolution path should still be online to be more promising.

— Guest
3:19:23

DeepSeek's papers read like school papers but use industry resources

The guest evaluates DeepSeek's paper style: very much like a school paper, low-key, not flashy, not promoting everywhere, but schools generally can't do such large models; they write school-style things with company industry-level resources, very different from other industry players. He adds this may relate to the team model — they incubated themselves, naturally a somewhat special team.

— Guest

In their own words · checked verbatim

But for a long time before, not just us — I think this community, including DeepSeek itself, and today's related papers will cover this too, including DeepSeek itself — actually everyone previously默认 we need another model, commonly called a reward model, to help judge whether my model's generation is correct, and then use this set to do reinforcement learning.

但是之前在很长一段时间 不光是我们 就我觉得这个community 包括Deepseek自己 今天也会讲到相关的paper 包括Deepseek自己 其实之前大家都是默认 我们需要另一个模型 我们俗称奖励模型 来帮助判断我的模型生成对不对 然后用这一套来做强化学习

He Junxian6:06

So I think, also got misled by OpenAI to some extent.

所以说我觉得就 也受到了一些open AI的误导吧

He Junxian7:06

But actually making this attempt is still very brave, because you have to spend a lot of compute at a very large scale, for example this kind of investment to explore something that, say, no one had really done before. MoE had been done by people before, for example bug expert, a small number of experts without shared, had been done and the effect was okay, so you could just follow that; the simplest thing might also be lower risk. Why insist on doing something different yourself?

但是其实要做这个尝试还是很勇敢 对因为你要在一个很大的规模上 花很大的算力 比如说这个投入去探索这种 你比如说以前大家都没有怎么做 大家MOE以前有人已经做过了 比如说就是bug expert 就少量 expert也不要shared 做过了就效果也还行 那你就直接照着做 其实最简单的也可能风险也比较低 为啥非要自己去搞一些不一样的东西

He Junxian58:57

Their train was very stable, the whole train, they didn't experience any loss spike.

就是他们的train非常的稳定 就是整个train 他们都没有经历任何的这个 loss spike

narration1:42:25

But almost no one had done this relatively successfully in real large-scale training.

但是几乎没有人在真的大规模上训练上 做过比较成功的做过这个事情

narration1:53:31

the improvement is attributed to boosting the correct response

the improvement is attributed to boosting the correct response

narration2:40:56

Then they found that after not using a reward model, this thing worked easily.

然后发现不用讲理模型之后 这个东西就很容易就work了

Guest3:19:23

I think their papers are very much like school, very much like school papers. Their resources are of course many; schools generally can't do such large models. But they write these things with company industry-level resources.

我觉得他们的paper很像学校的 很像学校的paper 他们资源当然很多了 学校一般做不了这么大的模型 但是他们就是拿着公司业界的 这种level的资源 写的这些东西

Guest3:19:23

Figures

High-Flyer A100 card count50008:07
DeepSeek MoE proprietary expert count641:01:57
MLA compression of KVCache93% reduction1:20:06
DeepSeek V3 training token count14.8T1:41:25
DeepSeek V3 training card count2000 H8001:43:28
Process-supervision reward model PPO debugging durationtwo or three months3:18:23

Glossary

MoE / Mixture of Experts
A model composed of multiple expert sub-networks, activating only some experts each time to reduce compute cost.
MLA / Multi-head Latent Attention
An attention mechanism proposed by DeepSeek that compresses K and V into a low-dimensional latent vector, greatly reducing KVCache.
MTP / Multi Token Prediction
Predicting multiple future tokens at once during training, increasing training signal density.
FP8 / 8-bit floating point
A low-precision training format that reduces memory and compute overhead but easily causes training instability.
PPO / Proximal Policy Optimization
A reinforcement learning algorithm often used for large-model alignment training.
rule-based reward
Using fixed rules (such as answer correctness, code compilation passing) as the reward signal, not relying on a reward model.

How to listen

Who it's for

AI engineers and researchers who want to understand DeepSeek's technical路线 and engineering decisions, plus founders watching large-model cost and efficiency.

Skip

Listeners uninterested in the math of MoE and MLA can skip the architecture section from 1:00:00 to 1:30:00.