TL;DR

A comparison of outcome supervision and process supervision for mathematical reasoning, centered on whether feedback should evaluate only the final answer or intermediate steps as well.

Why it matters

The granularity of a verifier changes both the learning signal and the failure modes. Process supervision can make reward more informative but is costlier to produce.

Key findings

  1. 01

    Step-level feedback can identify where reasoning goes wrong.

  2. 02

    Verifier design is a data and supervision decision, not merely a model choice.

Scaling dimensions

Models, methods & benchmarks

Algorithms
Process supervision

Topics

Read the original sourcearXiv