Tracking how reinforcement learning scales across models, compute, environments, rewards and agents — and helping teams turn frontier research into working systems.
A work-in-progress study of long-horizon agent reinforcement learning that uses privileged process supervision to redistribute verified trajectory-level credit across executable actions.
An empirical decomposition of LLM reinforcement-learning batch scaling into sample-indexed learning behavior and systems throughput, with a decision rule for when a larger batch lowers wall-clock time-to-target.
A controlled study of where reasoning diversity disappears during RLVR, distinguishing failure to initiate a solution family from failure to execute it after the first reasoning choice.
A compute-centric systematization of RL-style reasoning-model post-training, covering PPO and GRPO as well as intra-model, inter-model, synchronous, and asynchronous parallelism.
A controlled examination of R1-Zero-style reinforcement learning that separates robust training effects from claims that are sensitive to setup, reward design, and evaluation choices.
An official GLM-5.3 release note describing a post-training-only update on the same base model as GLM-5.2, with expanded executable environments for complex coding, security, and long-horizon professional workflows.
A technical report on Kimi K3, a 2.8T-parameter sparse multimodal model with 104B active parameters and a one-million-token context window, post-trained with reinforcement learning across general, agentic, coding, and reasoning domains.
An official release of the 750B-A40B GLM-5.2 model, combining a one-million-token context, IndexShare sparse-attention reuse, and larger-scale agentic reinforcement learning for long-horizon tasks.
An official release of MiniMax-M3, a natively multimodal sparse model with about 428B total and 23B active parameters, a one-million-token context window, and agent-oriented coding and cowork capabilities.
An official GLM-5.1 model update focused on keeping an agent productive across longer coding and engineering runs through repeated execution, inspection, diagnosis, and strategy revision.
A GLM-5 technical report centered on agentic engineering, combining a more efficient long-context architecture with asynchronous reinforcement-learning infrastructure and agent RL for complex, long-horizon software tasks.
A practical framework for locating the binding constraint in an RL system across environments, rollout generation, reward quality, optimization, and evaluation.
A decision-ready landscape of models, papers, competitors, and technical options for a specific strategic question.
02
AI Evaluation Sprint
Custom evaluation sets and a quality, cost, latency, and regression baseline for your model or agent system.
03
Agent Engineering Sprint
Workflow design, implementation, evaluation, observability, and handover for production-oriented AI agents.
06
RL Scaling Weekly
One useful briefing. Once a week.
The most important developments in RL scaling, post-training and agentic reinforcement learning — filtered and explained.
By subscribing, you agree to receive the weekly email. Unsubscribe at any time.
RL Scaling is an independent AI research and engineering studio. We publish source-linked research notes, explain why developments matter, and work with teams on frontier research, evaluation, and agent systems.