TL;DR

A concise reading map of weight-level and behavioral evidence on whether RLVR mainly reweights existing capabilities or creates new reasoning mechanisms.

Why it matters

The elicitation-versus-creation question affects how teams interpret pass@1 gains, solution coverage, weight updates, and the expected value of longer RL runs.

Key findings

  1. 01

    The post summarizes work suggesting RLVR updates can be sparse, low-rank, or less disruptive than supervised fine-tuning.

  2. 02

    At the behavioral level, it contrasts evidence that RLVR concentrates probability on existing solutions with prolonged-RL results that report expansion beyond sampled base-model behavior.

  3. 03

    The author labels the piece as informal synthesis and invites contradictory evidence; it should be read as a map of the debate.

Scaling dimensions

Models, methods & benchmarks

Algorithms
RLVR, SFT

Topics

Read the original sourceLessWrong