TL;DR

An empirical study of reinforcement learning directly from open base models, focused on making zero-RL experiments easier to reproduce and compare across model families.

Why it matters

Open reproduction is the fastest way to learn whether reasoning RL transfers beyond a small set of proprietary or heavily curated training stacks.

Key findings

  1. 01

    Base-model choice materially changes zero-RL behavior.

  2. 02

    Stable comparisons require aligned prompts, rewards, rollout budgets, and evaluation procedures.

Scaling dimensions

Models, methods & benchmarks

Algorithms
Zero RL

Topics

Read the original sourcearXiv