Independent AI research and engineering studio

RL ScalingResearch & Engineering
for Frontier AI

Tracking how reinforcement learning scales across models, compute, environments, rewards and agents — and helping teams turn frontier research into working systems.

01

Latest research

View all 47 items →
02

Frontier model reports

All model releases →
release

GLM-5.3: Frontier Coding with Emergent Cyber Capabilities

Z.ai

An official GLM-5.3 release note describing a post-training-only update on the same base model as GLM-5.2, with expanded executable environments for complex coding, security, and long-horizon professional workflows.

EnvironmentsAgentsRolloutsRewards
Read
release

Kimi K3: Open Frontier Intelligence

Moonshot AI

A technical report on Kimi K3, a 2.8T-parameter sparse multimodal model with 104B active parameters and a one-million-token context window, post-trained with reinforcement learning across general, agentic, coding, and reasoning domains.

ModelsComputeAgentsEnvironments
Read
release

GLM-5.2: Built for Long-Horizon Tasks

Z.ai

An official release of the 750B-A40B GLM-5.2 model, combining a one-million-token context, IndexShare sparse-attention reuse, and larger-scale agentic reinforcement learning for long-horizon tasks.

ModelsComputeAgentsEnvironments
Read
release

MiniMax-M3: Native Multimodal Intelligence at 1M Context

MiniMax

An official release of MiniMax-M3, a natively multimodal sparse model with about 428B total and 23B active parameters, a one-million-token context window, and agent-oriented coding and cowork capabilities.

ModelsComputeAgentsRollouts
Read
release

GLM-5.1: Towards Long-Horizon Tasks

Z.ai

An official GLM-5.1 model update focused on keeping an agent productive across longer coding and engineering runs through repeated execution, inspection, diagnosis, and strategy revision.

AgentsEnvironmentsRolloutsModels
Read
release

GLM-5: from Vibe Coding to Agentic Engineering

Z.ai

A GLM-5 technical report centered on agentic engineering, combining a more efficient long-context architecture with asynchronous reinforcement-learning infrastructure and agent RL for complex, long-horizon software tasks.

AgentsRolloutsComputeModels
Read
03

Featured analysis

04

Browse by topic

All topics →
05

Research into systems

Turn frontier AI research into working systems.

Focused engagements for teams facing a high-value research, evaluation, or engineering decision.

Discuss a project
01

Frontier AI Research Sprint

A decision-ready landscape of models, papers, competitors, and technical options for a specific strategic question.

02

AI Evaluation Sprint

Custom evaluation sets and a quality, cost, latency, and regression baseline for your model or agent system.

03

Agent Engineering Sprint

Workflow design, implementation, evaluation, observability, and handover for production-oriented AI agents.

06

RL Scaling Weekly

One useful briefing.
Once a week.

The most important developments in RL scaling, post-training and agentic reinforcement learning — filtered and explained.

RL Scaling is an independent AI research and engineering studio. We publish source-linked research notes, explain why developments matter, and work with teams on frontier research, evaluation, and agent systems.

About the studio →