Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Kilo Gateway logo

Model details

GLM-5.1

GLM-5.1 is positioned as Z.ai's next-generation flagship for agentic software engineering, designed to remain productive across unusually long autonomous sessions. Rather than plateauing after initial gains like earlier models, it is built to break down ambiguous problems, run experiments, read results, and iterate its strategy over hundreds of rounds and thousands of tool calls, sustaining optimization far beyond first-pass attempts. This long-horizon orientation shapes both its evaluation profile and its practical deployment, making it well suited to multi-step engineering workflows where persistence and judgment matter as much as raw capability.

The model's coding-focused design is reflected in its benchmark performance, where it achieves state-of-the-art results on SWE-Bench Pro and leads its predecessor by a wide margin on NL2Repo repository generation and Terminal-Bench 2.0 real-world terminal tasks. Z.ai highlights qualitative strengths in handling complex software engineering challenges, including cybersecurity-oriented work and the ability to rewrite CUDA kernels. As an open-weight release, GLM-5.1 gives teams the flexibility to self-host while still benefiting from a model explicitly tuned for autonomous, tool-driven development loops rather than short, single-shot completions.

Kilo Gatewayz-ai/glm-5.1glm

Quick Info

Powered by
Provider
Kilo Gateway
Model key
z-ai/glm-5.1
Release date
Apr 7, 2026
Last updated
Apr 7, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$1.40
Output token cost
$4.40

Limits

Output tokens
131,072 tokens
Context window
202,752 tokens

Transparent token rates

Compare GLM-5.1 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GLM-5.1

Ofox

Official sourceAnnouncement

Z.ai's GLM-5.2 announcement dated June 16, 2026 explicitly references GLM-5.1 as its predecessor and characterizes the upgrade as a substantial leap in long-horizon task capability, including the first solid 1M-token context. The post lists specific GLM-5.1-relative improvements in GLM-5.2, such as the IndexShare spars For GLM-5.1 specifically, the page establishes that GLM-5.1 lacked a reliable 1M context and that GLM-5.2's expanded 1M-context training targets coding-agent scenarios including large-scale implementation, automated research, performance optimization, and complex debugging. This makes the GLM-5.2 blog a credible, model

Kilo Gateway

CoverageLeaks

GLM-5.1, Z.ai's post-training upgrade to the GLM-5 foundation model, was released as open weights under the MIT license (available at huggingface.co/zai-org/GLM-5.1), making it one of the most permissively licensed large models in the current landscape. The architecture remains a 744 billion-parameter Mixture-of-Expert According to the report, GLM-5.1 claims the top score on SWE-Bench Pro at 58.4, surpassing GPT-5.4 (57.7) and Claude Opus 4.6 (57.3), and was trained entirely on Huawei Ascend 910B chips with no Nvidia hardware involvement — a notable milestone for non-Western AI compute infrastructure. At 744B parameters, the model is

OrcaRouter

CoverageBenchmark

Z.ai released GLM-5.1 on April 7, 2026 as an open-weight mixture-of-experts coding model under the MIT license, available on Hugging Face (zai-org/GLM-5.1). According to the source, the model has 754B total parameters with 40B active per token, a 200K-token context window, up to 128K-131K max output, and uses DSA spars The source reports that at launch GLM-5.1 achieved the top open-weights score on SWE-bench Pro at 58.4, ahead of GPT-5.4 (57.7) and Claude Opus 4.6 (57.3), and scored 63.5 on Terminal-Bench 2.0. It is compatible with transformers, vLLM, SGLang, KTransformers, and xLLM runtimes, and integrates with Claude Code and Cline

Kilo Gateway

Coverage

Z.AI Releases GLM-5.1: A Next-Generation Agentic Model Built for Long-Horizon Engineering Tasks and Rewrites CUDA Kernels

Kilo Gateway

Coverage

Z.ai raises prices for its most advanced AI model, GLM-5.1, by at least 8% compared to GLM-5 Turbo, joining Alibaba and Tencent as demand for agentic AI surges. This signals vendor-level price pressure for advanced models and could increase deployment and inference costs for developers and enterprises.

Kilo Gateway

Coverage

Price increases follow prior 30% hike as company targets monetization despite widening losses and strong demand for agentic AI services

CrossModel

CoverageBenchmark

InferenceX provides a technical overview distinguishing GLM-5 from its follow-up GLM-5.1. GLM-5 scales from 355B parameters (32B active) in GLM-4.5 to 744B parameters (40B active) with 28.5T pre-training tokens, released February 11, 2026 under MIT license. GLM-5.1 is described as the follow-up point release achieving Both models are served through the Z.ai API with a 200K context window and 128K maximum output. The technical deep-dive details that GLM-5 integrates DeepSeek Sparse Attention (DSA) to reduce deployment cost while preserving long-context capacity, alongside an asynchronous RL infrastructure called "slime" that decouple

CrossModel

Coverage

NYU Shanghai's RITS library published a third-party recap describing GLM-5.1 as a 754-billion-parameter open-weight Mixture-of-Experts model released by Z.ai on April 7, 2026, licensed under MIT and designed for agentic engineering. The recap reports that GLM-5.1 reached the #1 position on SWE-Bench Pro at 58.4%, ahead The write-up frames GLM-5.1 as a Dynamic Sparse Attention MoE with roughly 40 billion active parameters per token, capable of autonomously sustaining coding tasks for up to eight hours across hundreds of iterations. It cites demonstrations including a complete Linux desktop system built over an eight-hour, 655-iteratio

CrossModel

Official sourceRelease Notes

Z.ai's official developer documentation release notes confirm GLM-5.1 was released on April 7, 2026 as a model designed for long-horizon tasks capable of working independently for up to 8 hours in a single run. The release notes describe it as enabling a full loop from planning and execution to iterative refinement and According to the official docs, GLM-5.1 achieves comprehensive capability alignment with Claude Opus 4.6 and was built with multi-turn SFT, RL, and a process-based training approach. The documentation contextualizes GLM-5.1 within the broader model lineup, showing subsequent releases including GLM-5.2 (June 16, 2026, w

Videos about GLM-5.1

More models around GLM-5.1