TL;DR

DeepSeek-R1 documents a reasoning-model training pipeline centered on reinforcement learning, including a zero-style RL stage and a staged route that combines supervised data with RL.

Why it matters

The report made RL for verifiable reasoning a mainstream scaling path and created a concrete reference point for open reproduction efforts.

Key findings

  1. 01

    Verifiable rewards can support substantial reasoning-oriented post-training.

  2. 02

    A staged training recipe can balance emergent reasoning with usability and language quality.

Scaling dimensions

Models, methods & benchmarks

Models
DeepSeek-R1, DeepSeek-R1-Zero
Algorithms
Reinforcement learning, GRPO

Topics

Read the original sourcearXiv