Research index

What is actually
moving the frontier?

Source-linked research notes with original summaries and a clear view of why each development matters for scaling intelligent systems.

28 items
analysis

The RL Scaling Bottleneck Map

RL Scaling

A practical framework for locating the binding constraint in an RL system across environments, rollout generation, reward quality, optimization, and evaluation.

ComputeDataEnvironmentsRollouts
Read
release

GLM-5.3: Frontier Coding with Emergent Cyber Capabilities

Z.ai

An official GLM-5.3 release note describing a post-training-only update on the same base model as GLM-5.2, with expanded executable environments for complex coding, security, and long-horizon professional workflows.

EnvironmentsAgentsRolloutsRewards
Read
release

GR2 Technical Report

Yufei Li

A production-oriented technical report on a generative reasoning re-ranker trained with semantic-ID mid-training, teacher-trace distillation, on-policy distillation, and reinforcement learning from verifiable ranking rewards.

RewardsDataComputeEnvironments
Read
release

GLM-5.2: Built for Long-Horizon Tasks

Z.ai

An official release of the 750B-A40B GLM-5.2 model, combining a one-million-token context, IndexShare sparse-attention reuse, and larger-scale agentic reinforcement learning for long-horizon tasks.

ModelsComputeAgentsEnvironments
Read
release

MiniMax-M3: Native Multimodal Intelligence at 1M Context

MiniMax

An official release of MiniMax-M3, a natively multimodal sparse model with about 428B total and 23B active parameters, a one-million-token context window, and agent-oriented coding and cowork capabilities.

ModelsComputeAgentsRollouts
Read
paper

Self-Rewarding Language Models

Meta FAIR

This work studies language models that generate candidate responses and also provide the preference signal used to improve subsequent iterations.

RewardsDataModels
Read
paper

Let's Verify Step by Step

OpenAI

A comparison of outcome supervision and process supervision for mathematical reasoning, centered on whether feedback should evaluate only the final answer or intermediate steps as well.

RewardsDataModels
Read