TL;DR

An official GLM-5.1 model update focused on keeping an agent productive across longer coding and engineering runs through repeated execution, inspection, diagnosis, and strategy revision.

Why it matters

Long-horizon scaling is not only about accepting more tokens. GLM-5.1 focuses on whether a model can continue finding useful improvements after hundreds of iterations and thousands of tool calls instead of plateauing after the first attempt.

Key findings

  1. 01

    Z.ai positions GLM-5.1 as a long-horizon agentic-engineering update rather than a separate base-model technical report.

  2. 02

    The release says the model can sustain optimization over hundreds of rounds and thousands of tool calls by revisiting results and revising its strategy.

  3. 03

    The authors report gains over GLM-5 on SWE-Bench Pro, NL2Repo, Terminal-Bench 2.0, and CyberGym; these are vendor-reported results and depend on the stated harnesses and evaluation settings.

Scaling dimensions

Models, methods & benchmarks

Models
GLM-5.1, GLM-5
Algorithms
Long-horizon agent post-training, Iterative tool-use optimization
Benchmarks
SWE-Bench Pro, NL2Repo, Terminal-Bench 2.0, CyberGym

Topics

Read the original sourceZ.ai release