TL;DR
DeepSeek-R1 documents a reasoning-model training pipeline centered on reinforcement learning, including a zero-style RL stage and a staged route that combines supervised data with RL.
Why it matters
The report made RL for verifiable reasoning a mainstream scaling path and created a concrete reference point for open reproduction efforts.
Key findings
- 01
Verifiable rewards can support substantial reasoning-oriented post-training.
- 02
A staged training recipe can balance emergent reasoning with usability and language quality.
Scaling dimensions
Models, methods & benchmarks
- Models
- DeepSeek-R1, DeepSeek-R1-Zero
- Algorithms
- Reinforcement learning, GRPO
Topics