Research index

What is actually
moving the frontier?

Source-linked research notes with original summaries and a clear view of why each development matters for scaling intelligent systems.

29 items
release

Kimi K3: Open Frontier Intelligence

Moonshot AI

A technical report on Kimi K3, a 2.8T-parameter sparse multimodal model with 104B active parameters and a one-million-token context window, post-trained with reinforcement learning across general, agentic, coding, and reasoning domains.

ModelsComputeAgentsEnvironments
Read
release

Sakana Fugu Technical Report

Sakana AI

A technical report on Fugu and Fugu-Ultra, language-model orchestrators trained to construct query-adaptive scaffolds for teams of heterogeneous LLM agents.

AgentsComputeRolloutsModels
Read
release

GLM-5.2: Built for Long-Horizon Tasks

Z.ai

An official release of the 750B-A40B GLM-5.2 model, combining a one-million-token context, IndexShare sparse-attention reuse, and larger-scale agentic reinforcement learning for long-horizon tasks.

ModelsComputeAgentsEnvironments
Read
release

MiniMax-M3: Native Multimodal Intelligence at 1M Context

MiniMax

An official release of MiniMax-M3, a natively multimodal sparse model with about 428B total and 23B active parameters, a one-million-token context window, and agent-oriented coding and cowork capabilities.

ModelsComputeAgentsRollouts
Read
blog

Some Interesting Papers on RLVR

CarolusRenniusVitellius

A concise reading map of weight-level and behavioral evidence on whether RLVR mainly reweights existing capabilities or creates new reasoning mechanisms.

ModelsRolloutsData
Read
release

GLM-5.1: Towards Long-Horizon Tasks

Z.ai

An official GLM-5.1 model update focused on keeping an agent productive across longer coding and engineering runs through repeated execution, inspection, diagnosis, and strategy revision.

AgentsEnvironmentsRolloutsModels
Read
release

GLM-5: from Vibe Coding to Agentic Engineering

Z.ai

A GLM-5 technical report centered on agentic engineering, combining a more efficient long-context architecture with asynchronous reinforcement-learning infrastructure and agent RL for complex, long-horizon software tasks.

AgentsRolloutsComputeModels
Read
paper

Self-Rewarding Language Models

Meta FAIR

This work studies language models that generate candidate responses and also provide the preference signal used to improve subsequent iterations.

RewardsDataModels
Read
paper

Let's Verify Step by Step

OpenAI

A comparison of outcome supervision and process supervision for mathematical reasoning, centered on whether feedback should evaluate only the final answer or intermediate steps as well.

RewardsDataModels
Read
paper

Proximal Policy Optimization Algorithms

OpenAI

PPO introduces clipped and adaptive objectives intended to make policy-gradient optimization simpler to implement while limiting excessively large policy updates.

RolloutsComputeModels
Read