Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
302.AI logo

Model details

GLM-5.1

GLM-5.1 is designed as a next-generation agentic model centered on long-horizon engineering work, with a parameter scale around 754 billion that positions it among the largest open-weight language models available. The model's architecture is purpose-built for sustained autonomous execution, capable of independently planning, executing, and improving its own work over periods exceeding eight hours on a single task. Rather than being optimized for brief, discrete interactions, GLM-5.1 is engineered to deliver complete, engineering-grade results that require continuous reasoning and self-correction across extended timeframes. Its capabilities extend to demanding technical tasks such as rewriting CUDA kernels, and it achieves state-of-the-art performance on SWE-Bench Pro, a benchmark specifically measuring software engineering problem-solving ability.

The model represents a post-training refinement of GLM-5, developed with reinforced learning specifically targeting coding performance. This targeted post-training investment yielded approximately 28% improvement in coding capabilities over its predecessor while maintaining strengths in agentic engineering workflows. The combination of large-scale pre-training with specialized post-training for software engineering creates a model particularly suited for autonomous coding agents, integrated development environments, and continuous engineering tasks. Its availability across platforms including Ollama, where it has garnered millions of downloads, reflects strong community adoption for local deployment scenarios, while cloud providers offer API access for enterprise integration. The emphasis on agentic behavior—self-directed task completion, adaptive problem-solving, and extended operational duration—makes GLM-5.1 a fit for development workflows requiring persistent AI assistance beyond traditional chat-based interactions.

302.AIglm-5.1glm

Quick Info

Powered by
Provider
302.AI
Model key
glm-5.1
Release date
Apr 7, 2026
Last updated
Apr 7, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$1.40
Output token cost
$4.40

Limits

Output tokens
131,072 tokens
Context window
200,000 tokens

Transparent token rates

Compare GLM-5.1 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GLM-5.1

302.AI

CoverageAnalysis

The Baidu Cloud technical analysis (dated Sep 7, 2026) positions GLM-5.1 as a new-generation open-source flagship engineered for complex multi-turn interactions and long-sequence reasoning. The supplied excerpt describes a Mixture-of-Experts architecture that maintains 100B+ parameter scale while activating only 40B pa The article provides architectural detail not fully covered elsewhere: the MoE uses 256 routing experts plus 1 shared expert with a Top-8 activation strategy and dynamic routing, with input tokens first processed through 3 dense layers before a gating network computes expert weights and the top 8 experts run in paralle

302.AI

CoverageBenchmark

GLM-5.1 is Z.ai's open-weight mixture-of-experts coding model released on April 7, 2026 under the MIT license, with 744 billion total parameters (40 billion active per token), a 200K-token context window, and up to 128K–131K maximum output. It introduces DSA sparse attention and is positioned as a coding- and agent-fir At launch, GLM-5.1 took the top published SWE-bench Pro score at 58.4, ahead of GPT-5.4 (57.7) and Claude Opus 4.6 (57.3), and scored 63.5 on Terminal-Bench 2.0, establishing it as the leading open-weight coding model at release. The MorphLLM piece flags a "glm47 tool-call parser gotcha" relevant to integrators and not

302.AI

CoverageBenchmark

Zhipu AI (branded internationally as Z.ai) released GLM-5.1 on March 27, 2026 as an incremental upgrade to its GLM-5 flagship, claiming a self-reported coding benchmark score of 45.3 versus Claude Opus 4.6's 47.9 when measured using the Claude Code evaluation harness — approximately 94.6% of Opus 4.6 and a 28% gain ove For developers accessing GLM-5.1 via the 302.AI API endpoint at api.302.ai/v1, pricing is positioned aggressively below Western frontier tiers: the standalone GLM-5 API is listed at $1.00 per million input tokens and $3.20 per million output tokens, while Z.ai's GLM Coding Plan is offered at a promotional $3/month for

302.AI

CoverageBenchmark

The InferenceX architectural overview identifies GLM-5.1 as Z.ai's follow-up point release on the GLM-5 architecture, described as a next-generation flagship for agentic engineering with significantly stronger coding capabilities than its predecessor. The supplied excerpt cites the GLM-5.1 model card for state-of-the-a The overview also establishes shared GLM-5/5.1 architectural context: 744B parameters (40B active), 28.5T pre-training tokens, DeepSeek Sparse Attention to reduce deployment cost while preserving long-context capacity, and an asynchronous RL infrastructure named "slime" that decouples generation from training, as descr

302.AI

CoverageRelease Notes

The Z.ai developer documentation release notes list GLM-5.1 with a release date of 2026-04-07, describing it as designed for long-horizon tasks where it can work independently for up to 8 hours in a single run, enabling a full loop from planning and execution to iterative refinement and final delivery. The page states The same official index page also lists subsequent GLM releases, providing useful context for the GLM family trajectory: GLM-5.2 (2026-06-16) with 1M lossless context; GLM-5.3 (2026-08-18) with stronger coding and emergent cybersecurity capabilities; and GLM-5.3-Flash (2026-08-26) with native visual capabilities and a

Videos about GLM-5.1

More models around GLM-5.1