GLM-5.1 is a 754-billion parameter flagship model built specifically for agentic engineering—tasks that demand sustained reasoning and autonomous execution over hours rather than minutes. Its architecture is engineered to stay effective across long horizons, unlike predecessors that plateau after initial gains. The model handles ambiguous problems with adaptive judgment, breaks complex issues into experiments, reads results, and revises strategy through repeated iteration. This design makes it particularly strong at automating multi-hour engineering workflows, from writing and rewriting CUDA kernels to executing complex software development cycles end-to-end.
While explicit pre-training details remain limited in public sources, GLM-5.1 represents a significant leap over GLM-5, achieving state-of-the-art performance on SWE-Bench Pro with a score of 58.4—outpacing GPT-5.4, Opus 4.6, and Gemini 3.1 Pro on complex software engineering benchmarks. The model has been released as open-weight through Zhipu's GitHub and HuggingFace repositories, enabling community access and further development. Practical applications include automated repo generation, real-world terminal task execution, and cybersecurity problem-solving, with the model sustaining productive performance across extended sessions in ways previous GLM generations could not.