TL;DR
A community argument that RLVR is operationally important but that most language-model knowledge and available reasoning strategies still originate in pretraining and supervised imitation.
Why it matters
The source of capability changes what teams should expect from additional RL. If RL mostly improves strategy selection over existing mechanisms, environment quality and sampling may matter more than assuming every longer run creates new skills.
Key findings
- 01
The post argues that pretraining and SFT supply most of the model's knowledge and reasoning repertoire while RLVR sharpens when and how to use it.
- 02
It reviews evidence that base models can sometimes reach RLVR solutions through enough sampling, alongside counterevidence from prolonged RL that reports genuinely new solution paths.
- 03
The author explicitly presents this as a current mental model rather than a settled conclusion.
Scaling dimensions
Models, methods & benchmarks
- Algorithms
- RLVR, Supervised fine-tuning, Pretraining
Topics