R1's most valuable output isn't the model, it's the aha moment
DeepSeek was the first to solve the riddle O1 posed: without being taught, the model learns reflection and error correction on its own through reinforcement learning. R1-Zero is cleaner than R1, but R1 is the one that works.
The video won't play here. Listen to the audio instead:
The argument · tap a timestamp to hear it
The aha moment wasn't written into the model by humans
The R1 paper uses the word incentivizing in its title, behind which is OpenAI researcher Hongi's talk at MIT, ‘Don't teach, incentivize’—don't teach the model how to do it, but tell it what's good and what's bad, and let it figure it out itself. The self-correction in the paper, ‘Wait, wait, there's an aha moment I can fly here’, wasn't written into the model by the researchers; it emerged on its own during reinforcement learning. This is like the divine move in Go—a behavioral pattern human developers never thought of.
— Pan JiayiPost-training compute is just a drizzle compared to pre-training
DeepSeek V3 pre-training cost 5.3M USD, mid-training used 0.2M to extend the context from 8k to 128k, and post-training cost only 0.01M—less than 0.2% of the pre-training cost. This number shows that post-training and reinforcement learning scaling is still in a very early stage, with an extremely high return on investment. For the entire R1 training, if optimized well, it might cost just 100k USD; at the absolute most, 1M USD.
— Pan JiayiGRPO drops the value function because it's inaccurate
PPO requires training an additional value model to finely judge whether each step is right or wrong, but in long chain-of-thought scenarios this assumption doesn't hold: the supervision signal only exists at the last step, the sequence is thousands of tokens long, and the model has learned to correct itself—whether the 3 in ‘1+1=3’ is right or wrong becomes unclear, because it can fix it later. The value function cannot be precise, and it consumes enormous compute, so DeepSeek simply discards it and only does policy gradients. This is the starting point of GRPO.
— Pan JiayiThe reward function uses only rules, not neural networks
R1-Zero's reward has two parts: accuracy reward checks whether the answer is correct, format reward checks whether the format is correct—requiring the model to put the thinking process in a think tag and the answer in an answer tag. They deliberately avoid process reward models or outcome reward models, because a neural network reward model will be reward hacked: the model will find loopholes in the reward function, for example discovering that outputting many emojis gets a high score, and then output emojis like crazy. Using ground truth reward avoids any possibility of reward hacking.
— Pan JiayiThe performance gain comes from output getting 20 times longer
Figure 2 shows R1's accuracy on AIME rising from the teens all the way to around 70%, and after pausing for 10,000 steps it roughly reaches O1's range. But Figure 3 explains why: during reinforcement learning the model discovers on its own that if it thinks more at inference time, generates more tokens, and uses more compute, performance improves. At step 0 the output is only a few hundred tokens; after 8,000 steps the average response length reaches about 10,000 tokens, an increase of nearly 20 times. This is the source of inference-time scaling.
— Pan JiayiDistillation beats small models doing their own reinforcement learning
DeepSeek open-sourced a series of distilled models from 1.5B to 70B in one go. The 1.5B distilled model already beats GPT-4o or Claude Sonnet by 10 to 20 percentage points on AIME. More critically, the comparison in Table 6: distilling R1 into Qwen 32B far exceeds Qwen 32B doing reinforcement learning itself in the R1-Zero way. The reason is that large models are more likely to explore complex and beneficial reasoning patterns, and handing that exploration to a small model is much better than having the small model explore from scratch.
— Pan JiayiBoth PRM and MCTS paths are dead ends
DeepSeek publicly disclosed its failed attempts. Problems with process reward models: on general tasks it's hard to define which step is the first and which is the second; evaluating whether the current step is correct is also difficult, and becomes even more ambiguous once the model can correct itself; automatic training results are unsatisfactory, and manual annotation is too expensive and doesn't scale; and as long as there is a reward model, it will inevitably bring reward hacking. Problems with MCTS: the search space at each step of a language model is about 10,000 words, the tree can have 10,000 layers, the search space is too large and the cost too high; the value function is unstable and hard to estimate. AlphaGo's success is hard to reproduce in general language model reasoning scenarios.
— Pan JiayiR1 training cost might be only 100k USD
Rough estimate: the entire training used about 10,000 steps of reinforcement learning, each step about 1,000 responses, each response at most 10,000 tokens, generating about 100 billion tokens in total. At R1's API price of 2.2 USD per million tokens, these tokens cost just over 200k USD. Considering that DeepSeek has a large margin, and ignoring part of the model training overhead, the order of magnitude should be about right. The conclusion is that if optimized well enough it might cost just 100k USD, and at the absolute most 1M USD.
— Pan JiayiChain-of-thought reward models cut error rate from 15% to 1.5%
K1.5 tried two methods for math rewards. The traditional reward model connects the model's activations to a small MLP to directly predict right or wrong, with accuracy around 84.4%. The chain-of-thought reward model takes the question, the standard answer, and the model's answer together as input, and lets the reward model output a chain of thought to reason about right or wrong, fine-tuned with 800k labeled samples. Accuracy directly reaches 98.5%, and the error rate drops from 15% to 1.5%. This is not in the DeepSeek report, and everyone should move in this direction going forward.
— Pan JiayiWithout negative gradients, reinforcement learning can't get off the ground
K1.5 compared a simpler REST algorithm: sample 100 times, pick out the correct ones for fine-tuning, done. Its difference from GRPO is: it doesn't require staying close enough to the original policy, and it doesn't tell the model not to learn bad things—missing the negative gradient step. The result is that REST's performance is considerably worse, the entire test time scaling completely fails to take off, maybe only rising three or four points; while using K1.5 or R1's reinforcement learning algorithm can rise ten or twenty points, a world of difference.
— Pan JiayiIn their own words · checked verbatim
This is actually the only method we currently know of that can achieve superhuman performance.
这其实是我们现在 唯一已知的就是可以达到就是 superhuman performance的一个方法
Pan Jiayi1:13:14
A large model, compared to a small model, has better performance and is more likely to explore those complex, or some powerful, more beneficial, more complex [patterns].
大模型它想象比小模型来说 它的就是它的性能更好 它更有可能能探索到那些复杂的 或者说一些就是有力 更加有益的 更加复杂
Pan Jiayi1:32:23
So elegant algorithms or elegant techniques are often the simplest and cleanest techniques, right? Yes, yes.
所以优美的算法或者优美的技术 往往是最简单干净的技术是吗 是的 是的
Pan Jiayi2:39:00
Figures
| DeepSeek V3 pre-training cost | 5.3M USD | 38:48 |
| DeepSeek V3 post-training cost | 0.01M USD | 38:48 |
| R1 API price | 2.2 USD per million tokens | 1:44:30 |
| R1 training cost estimate | 100k to 1M USD | 1:45:30 |
| R1 output length growth | from a few hundred tokens to about 10,000 tokens, an increase of nearly 20 times | 1:09:08 |
| Traditional reward model accuracy | 84.4% | 2:01:40 |
| Chain-of-thought reward model accuracy | 98.5% | 2:01:40 |
| Chain-of-thought reward model training samples | 800k | 2:00:40 |
| R1 distillation data volume | reasoning 600k + non-reasoning 200k = 800k | 1:22:17 |
Glossary
- GRPO / Group Relative Policy Optimization
- A PPO variant proposed by DeepSeek that drops the value function, only does policy gradients, and normalizes using relative goodness within a group.
- reward hacking
- The model finds loopholes in the reward function and behaves in ways that get high scores but that humans don't want, such as outputting emojis like crazy.
- PRM / Process Reward Model
- Scores each step of reasoning rather than only the final result, but on general tasks it's hard to define steps and it's easily hacked.
- MCTS / Monte Carlo Tree Search
- Succeeded in AlphaGo, but in language model reasoning the search space is too large and the value function is hard to estimate.
- distillation
- Fine-tuning a small model with data generated by a large model, giving the small model reasoning ability, with better results than the small model doing its own reinforcement learning.
- curriculum sampling
- A K1.5 technique that gives the model problems whose difficulty matches its ability, letting it learn faster and better.
How to listen
Engineers and researchers who want to understand the technical details of R1 and K1.5, especially those working on post-training and reinforcement learning.
The first 15 minutes of background setup and the chit-chat at the end can be skipped; the core content starts at 35 minutes.