An RL environment defines the tasks, observations, actions, transitions, and outcomes through which a policy gains experience. For frontier AI, environment engineering increasingly includes browsers, code repositories, simulators, synthetic users, and tool APIs. Environment diversity and fidelity often determine whether gains transfer beyond a benchmark.
An official GLM-5.3 release note describing a post-training-only update on the same base model as GLM-5.2, with expanded executable environments for complex coding, security, and long-horizon professional workflows.
A LessWrong proposal to let policies report exploitable RLVR environment bugs after a rollout, reward high-quality verified reports, and use them to patch the training environment.
A technical report on Kimi K3, a 2.8T-parameter sparse multimodal model with 104B active parameters and a one-million-token context window, post-trained with reinforcement learning across general, agentic, coding, and reasoning domains.
A technical report on Athena-Brain-8B, an on-device embodied model trained through general supervised fine-tuning, general reinforcement learning, embodied-expert training, and model merging.
An official release of the 750B-A40B GLM-5.2 model, combining a one-million-token context, IndexShare sparse-attention reuse, and larger-scale agentic reinforcement learning for long-horizon tasks.
A model-family report on Ling-2.6 and Ring-2.6, combining architectural migration, long-context efficiency work, token-efficient reasoning objectives, and asynchronous agent reinforcement learning at trillion-parameter scale.
A technical report on the MiniMax-M2 family, pairing a 229.9B-parameter sparse MoE with agent-generated, verifiable trajectories and Forge, a scalable reinforcement-learning system for long-horizon agents.
An official GLM-5.1 model update focused on keeping an agent productive across longer coding and engineering runs through repeated execution, inspection, diagnosis, and strategy revision.
A technical report on Composer 2, a specialized coding model trained through continued pretraining followed by large-scale reinforcement learning on long-horizon software-engineering tasks.
A GLM-5 technical report centered on agentic engineering, combining a more efficient long-context architecture with asynchronous reinforcement-learning infrastructure and agent RL for complex, long-horizon software tasks.
WebArena provides self-hosted, realistic websites and benchmark tasks for agents that must navigate interfaces, maintain state, and complete multi-step goals.
Voyager combines an automatic curriculum, an executable skill library, and iterative prompting to build an open-ended Minecraft agent without model parameter updates.
This work frames task-distribution design as an optimization problem and introduces PAIRED, where an environment generator uses regret to create solvable but challenging curricula.