Topic

RLVR

Reinforcement Learning with Verifiable Rewards uses outcomes that can be checked by programs, tests, simulators, or formal rules. Math, code, and tool-use tasks are common examples. RLVR can make feedback cheaper and more scalable than human preference labeling, but only where the verifier captures the behavior that matters and resists reward gaming.

Key concepts

Verifiable outcomes

Programmatic graders

Reward hacking

Reasoning traces

Latest research

15 items
release

GR2 Technical Report

Yufei Li

A production-oriented technical report on a generative reasoning re-ranker trained with semantic-ID mid-training, teacher-trace distillation, on-policy distillation, and reinforcement learning from verifiable ranking rewards.

RewardsDataComputeEnvironments
Read
blog

Some Interesting Papers on RLVR

CarolusRenniusVitellius

A concise reading map of weight-level and behavioral evidence on whether RLVR mainly reweights existing capabilities or creates new reasoning mechanisms.

ModelsRolloutsData
Read
paper

Let's Verify Step by Step

OpenAI

A comparison of outcome supervision and process supervision for mathematical reasoning, centered on whether feedback should evaluate only the final answer or intermediate steps as well.

RewardsDataModels
Read