Research index

What is actually
moving the frontier?

Source-linked research notes with original summaries and a clear view of why each development matters for scaling intelligent systems.

21 items
analysis

The RL Scaling Bottleneck Map

RL Scaling

A practical framework for locating the binding constraint in an RL system across environments, rollout generation, reward quality, optimization, and evaluation.

ComputeDataEnvironmentsRollouts
Read
release

GR2 Technical Report

Yufei Li

A production-oriented technical report on a generative reasoning re-ranker trained with semantic-ID mid-training, teacher-trace distillation, on-policy distillation, and reinforcement learning from verifiable ranking rewards.

RewardsDataComputeEnvironments
Read
blog

Some Interesting Papers on RLVR

CarolusRenniusVitellius

A concise reading map of weight-level and behavioral evidence on whether RLVR mainly reweights existing capabilities or creates new reasoning mechanisms.

ModelsRolloutsData
Read
paper

Self-Rewarding Language Models

Meta FAIR

This work studies language models that generate candidate responses and also provide the preference signal used to improve subsequent iterations.

RewardsDataModels
Read
paper

Let's Verify Step by Step

OpenAI

A comparison of outcome supervision and process supervision for mathematical reasoning, centered on whether feedback should evaluate only the final answer or intermediate steps as well.

RewardsDataModels
Read