GLM-5.1 is a next-generation agentic model built around a 754 billion parameter architecture, engineered specifically for sustained software engineering and complex coding tasks. Unlike models that deliver quick wins and then plateau, GLM-5.1 is designed to stay productive over long horizons—it decomposes ambiguous problems, runs experiments, reads results, and identifies blockers with deliberate judgment. The model revisits its own reasoning, revises strategy through repeated iteration, and sustains optimization across hundreds of rounds and thousands of tool calls. This architecture-first focus on extended agentic execution sets it apart for tasks requiring deep persistence, like multi-hour autonomous coding sessions or rewriting CUDA kernels from first principles.
GLM-5.1 builds on its GLM-5 predecessor with significant improvements in coding capability, achieving state-of-the-art performance on SWE-Bench Pro and leading on benchmarks like NL2Repo repo generation and Terminal-Bench 2.0 for real-world terminal tasks. Sources indicate it was developed as a flagship agentic engineering model with coding at its core, matching the performance level of Opus 4.6 on programming tasks. Available as an open-weight release through HuggingFace, the model offers developers a large-scale foundation for building autonomous agents that can handle extended, multi-step engineering workflows with adaptive problem-solving and persistent execution.