TL;DR
An official GLM-5.1 model update focused on keeping an agent productive across longer coding and engineering runs through repeated execution, inspection, diagnosis, and strategy revision.
Why it matters
Long-horizon scaling is not only about accepting more tokens. GLM-5.1 focuses on whether a model can continue finding useful improvements after hundreds of iterations and thousands of tool calls instead of plateauing after the first attempt.
Key findings
- 01
Z.ai positions GLM-5.1 as a long-horizon agentic-engineering update rather than a separate base-model technical report.
- 02
The release says the model can sustain optimization over hundreds of rounds and thousands of tool calls by revisiting results and revising its strategy.
- 03
The authors report gains over GLM-5 on SWE-Bench Pro, NL2Repo, Terminal-Bench 2.0, and CyberGym; these are vendor-reported results and depend on the stated harnesses and evaluation settings.
Scaling dimensions
Models, methods & benchmarks
- Models
- GLM-5.1, GLM-5
- Algorithms
- Long-horizon agent post-training, Iterative tool-use optimization
- Benchmarks
- SWE-Bench Pro, NL2Repo, Terminal-Bench 2.0, CyberGym
Topics