TL;DR
DAPO presents an open reinforcement-learning system and a collection of training interventions aimed at improving large-scale reasoning optimization.
Why it matters
It moves discussion from a single optimizer name to the system-level choices—sampling, clipping, length handling, and reward stability—that often determine whether RL runs work.
Key findings
- 01
Large-scale RL performance depends on a bundle of interacting system and optimization choices.
- 02
Open recipes make ablation and independent verification more practical.
Scaling dimensions
Models, methods & benchmarks
- Algorithms
- DAPO, Policy optimization
Topics