Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
ZenMux logo

Model details

GLM 5.1

GLM-5.1 is a next-generation agentic model built around a 754 billion parameter architecture, engineered specifically for sustained software engineering and complex coding tasks. Unlike models that deliver quick wins and then plateau, GLM-5.1 is designed to stay productive over long horizons—it decomposes ambiguous problems, runs experiments, reads results, and identifies blockers with deliberate judgment. The model revisits its own reasoning, revises strategy through repeated iteration, and sustains optimization across hundreds of rounds and thousands of tool calls. This architecture-first focus on extended agentic execution sets it apart for tasks requiring deep persistence, like multi-hour autonomous coding sessions or rewriting CUDA kernels from first principles.

GLM-5.1 builds on its GLM-5 predecessor with significant improvements in coding capability, achieving state-of-the-art performance on SWE-Bench Pro and leading on benchmarks like NL2Repo repo generation and Terminal-Bench 2.0 for real-world terminal tasks. Sources indicate it was developed as a flagship agentic engineering model with coding at its core, matching the performance level of Opus 4.6 on programming tasks. Available as an open-weight release through HuggingFace, the model offers developers a large-scale foundation for building autonomous agents that can handle extended, multi-step engineering workflows with adaptive problem-solving and persistent execution.

ZenMuxz-ai/glm-5.1glm

Quick Info

Powered by
Provider
ZenMux
Model key
z-ai/glm-5.1
Release date
Apr 7, 2026
Last updated
Apr 7, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.8781
Output token cost
$3.5126

Limits

Output tokens
128,000 tokens
Context window
200,000 tokens

Transparent token rates

Compare GLM 5.1 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GLM 5.1

ZenMux

Official sourceOfficial

ZenMux lists the z-ai/glm-5.1 routing page as Active and provides per-provider gateway metadata for the model. The page shows three hosted providers—Streamlake (46.8% cache hit rate), SiliconFlow, and BigModel—each offering a 200K token context window. Streamlake's live metrics indicate roughly 37.2 tokens/second throu ZenMux's published pricing for z-ai/glm-5.1 on the routing page is $0.8781–1.1709 per million input tokens and $3.5126–4.098 per million output tokens across providers, with cache reads at $0.1903–0.2927 per million tokens. The page also surfaces a "Token Consumption" breakdown of public apps sending traffic to the mod

ZenMux

CoverageBenchmark

Morphllm's July 17, 2026 explainer synthesizes GLM-5.1 as a 754B-total mixture-of-experts coding model from Z.ai (formerly Zhipu AI) with 40B active parameters per token, a 200K-token context window, and 128K–131K max output, released April 7, 2026 under an MIT license. It reports the April-launch SWE-Bench Pro score o The piece details the model's DSA sparse attention architecture and explains that it is tuned for 8-hour autonomous agentic runs with hundreds of tool-call rounds. It positions GLM-5.1 as coding-first with open weights on zai-org/GLM-5.1, runs on transformers, vLLM, SGLang, KTransformers, and xLLM, and drops into Claud

ZenMux

Coverage

An NYU Shanghai RITS write-up corroborates Z.ai's April 7, 2026 release of GLM-5.1, describing it as a 754-billion-parameter open-weight MoE model (approximately 40 billion active parameters per token) built on Dynamic Sparse Attention, licensed under MIT and designed for agentic engineering. The model claims the #1 po The piece reports strong math and reasoning numbers (95.3% on AIME 2026, 86.2% on GPQA-Diamond, 83.8% on IMOAnswerBench) and highlights agentic engineering demos including an eight-hour, 655-iteration autonomous build of a complete Linux desktop system and a vector database throughput increase to 6.9× the initial versi

ZenMux

Official sourceRelease Notes

The Z.ai developer documentation release notes cover GLM-5.1 (April 7, 2026) alongside later releases. GLM-5.1 is described as designed for long-horizon tasks, capable of working independently for up to 8 hours in a single run, enabling a full loop from planning and execution to iterative refinement and final delivery, The release notes confirm GLM-5.1 achieves comprehensive capability alignment with Claude Opus 4.6 and was built with multi-turn SFT, RL, and a process-based training approach. The same page documents the later GLM-5.2 (June 16, 2026, 1M lossless context), GLM-5.3 (August 18, 2026, 50% gain over 5.2 on Z.ai Code Bench)

Videos about GLM 5.1

More models around GLM 5.1