Topic

Scaling Laws

Scaling laws describe empirical relationships between resources and performance. In reinforcement learning, useful laws must account for interaction data, reward quality, environment diversity, optimization pressure, and failure under proxy objectives—not only parameter count and training compute.

Key concepts

Compute-performance curves

Data scaling

Reward overoptimization

Emergent capability

Latest research

14 items
analysis

The RL Scaling Bottleneck Map

RL Scaling

A practical framework for locating the binding constraint in an RL system across environments, rollout generation, reward quality, optimization, and evaluation.

ComputeDataEnvironmentsRollouts
Read
release

Kimi K3: Open Frontier Intelligence

Moonshot AI

A technical report on Kimi K3, a 2.8T-parameter sparse multimodal model with 104B active parameters and a one-million-token context window, post-trained with reinforcement learning across general, agentic, coding, and reasoning domains.

ModelsComputeAgentsEnvironments
Read
release

GLM-5.2: Built for Long-Horizon Tasks

Z.ai

An official release of the 750B-A40B GLM-5.2 model, combining a one-million-token context, IndexShare sparse-attention reuse, and larger-scale agentic reinforcement learning for long-horizon tasks.

ModelsComputeAgentsEnvironments
Read
release

MiniMax-M3: Native Multimodal Intelligence at 1M Context

MiniMax

An official release of MiniMax-M3, a natively multimodal sparse model with about 428B total and 23B active parameters, a one-million-token context window, and agent-oriented coding and cowork capabilities.

ModelsComputeAgentsRollouts
Read
blog

Some Interesting Papers on RLVR

CarolusRenniusVitellius

A concise reading map of weight-level and behavioral evidence on whether RLVR mainly reweights existing capabilities or creates new reasoning mechanisms.

ModelsRolloutsData
Read