TL;DR
A comparison of outcome supervision and process supervision for mathematical reasoning, centered on whether feedback should evaluate only the final answer or intermediate steps as well.
Why it matters
The granularity of a verifier changes both the learning signal and the failure modes. Process supervision can make reward more informative but is costlier to produce.
Key findings
- 01
Step-level feedback can identify where reasoning goes wrong.
- 02
Verifier design is a data and supervision decision, not merely a model choice.
Scaling dimensions
Models, methods & benchmarks
- Algorithms
- Process supervision
Topics