TL;DR
A concise reading map of weight-level and behavioral evidence on whether RLVR mainly reweights existing capabilities or creates new reasoning mechanisms.
Why it matters
The elicitation-versus-creation question affects how teams interpret pass@1 gains, solution coverage, weight updates, and the expected value of longer RL runs.
Key findings
- 01
The post summarizes work suggesting RLVR updates can be sparse, low-rank, or less disruptive than supervised fine-tuning.
- 02
At the behavioral level, it contrasts evidence that RLVR concentrates probability on existing solutions with prolonged-RL results that report expansion beyond sampled base-model behavior.
- 03
The author labels the piece as informal synthesis and invites contradictory evidence; it should be read as a map of the debate.
Scaling dimensions
Models, methods & benchmarks
- Algorithms
- RLVR, SFT
Topics