Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Vercel AI Gateway logo

Model details

GLM 5.1

GLM 5.1 is a large-scale language model purpose-built for coding and long-running software engineering tasks, designed to operate as a self-directed agent rather than a quick turn-taking assistant. It uses a mixture-of-experts design with 754B total parameters and roughly 40B activated per pass, paired with DeepSeek Sparse Attention and a 200K context window that can stretch to 131K output tokens. That combination is meant to let the model keep a full engineering problem in working memory while it plans, edits, tests, and iterates, sustaining coherent behavior across frontend, backend, and systems work without constant human steering. The model also exposes a thinking mode, tool calling, and structured JSON output so it can plug into coding-agent frameworks as a drop-in reasoning engine.

GLM 5.1 sits on the GLM-5 base architecture but layers in a refined reinforcement-learning post-training pass that Z.ai highlights as the source of a 28% coding gain over its predecessor. The release emphasizes long-horizon execution, with Z.ai claiming the model can work independently on a single task for more than eight hours, planning and improving its own output along the way. On public evaluations it posts 77.8% on SWE-Bench Verified, 50.4% on HLE with tools, 92.7% on AIME 2025, 86.0% on GPQA Diamond, and scores 45.3 on Z.ai's internal coding suite, alongside a #1 placement on Vending Bench 2 among open-source models. Z.ai distributes the weights openly under an MIT-style license and frames GLM 5.1 as a flagship open alternative to closed frontier coding models, well suited for agentic engineering pipelines, research workflows, and production code generation where sustained autonomy and tool use matter more than single-turn chat.

Vercel AI Gatewayzai/glm-5.1glm

Quick Info

Powered by
Provider
Vercel AI Gateway
Model key
zai/glm-5.1
Release date
Apr 7, 2026
Last updated
Apr 7, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$1.40
Output token cost
$4.40

Limits

Output tokens
64,000 tokens
Context window
202,800 tokens

Transparent token rates

Compare GLM 5.1 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GLM 5.1

Vercel AI Gateway

Coverage

Zhipu is beginning to lean into monetization after a prolonged investment phase, raising prices for access to its most advanced AI models by at least 8% alongsi

Vercel AI Gateway

Coverage

The Chinese company said its new open-source model can continue to improve over hundreds of iterations, as AI vendors race to build tools that can handle longer software tasks.

Vercel AI Gateway

Official sourceRelease Notes

You can now access GLM 5.1 on Vercel's AI Gateway with no markup and no other provider accounts required.

Vercel AI Gateway

CoverageBenchmark

InferenceX's profile provides technical context tying GLM-5.1 to its GLM-5 base, noting GLM-5 scales to 744B total parameters with 40B active and 28.5T pre-training tokens under an MIT license, and that GLM-5.1 is a follow-up point release on the same architecture released April 7, 2026, also MIT-licensed. It cites the The profile confirms that both GLM-5 and GLM-5.1 are served through the Z.ai API with a 200K context window and 128K maximum output, and links the GLM-5.1 weight repository (created April 3, 2026) to Hugging Face. The excerpt makes no claims about Vercel AI Gateway hosting or pricing for GLM-5.1.

Vercel AI Gateway

Official sourceRelease Notes

The Z.AI release notes index logs a 2026-04-07 entry for GLM-5.1 stating the model is designed for long-horizon tasks and can work independently for up to 8 hours in a single run, covering planning, execution, iterative refinement, and final delivery. The release-notes excerpt credits GLM-5.1's gains to multi-turn SFT, The same release-notes page also lists subsequent versions—GLM-5.2 (2026-06-16) introducing a 1M lossless context, and GLM-5.3 / GLM-5.3-Flash (2026-08-26) adding native visual capabilities and stronger coding and cybersecurity results—confirming GLM-5.1's place as the April 2026 baseline within Z.AI's rapidly evolving

Videos about GLM 5.1

More models around GLM 5.1