Vercel AI Gateway
Zhipu is beginning to lean into monetization after a prolonged investment phase, raising prices for access to its most advanced AI models by at least 8% alongsi
Model details
GLM 5.1 is a large-scale language model purpose-built for coding and long-running software engineering tasks, designed to operate as a self-directed agent rather than a quick turn-taking assistant. It uses a mixture-of-experts design with 754B total parameters and roughly 40B activated per pass, paired with DeepSeek Sparse Attention and a 200K context window that can stretch to 131K output tokens. That combination is meant to let the model keep a full engineering problem in working memory while it plans, edits, tests, and iterates, sustaining coherent behavior across frontend, backend, and systems work without constant human steering. The model also exposes a thinking mode, tool calling, and structured JSON output so it can plug into coding-agent frameworks as a drop-in reasoning engine.
GLM 5.1 sits on the GLM-5 base architecture but layers in a refined reinforcement-learning post-training pass that Z.ai highlights as the source of a 28% coding gain over its predecessor. The release emphasizes long-horizon execution, with Z.ai claiming the model can work independently on a single task for more than eight hours, planning and improving its own output along the way. On public evaluations it posts 77.8% on SWE-Bench Verified, 50.4% on HLE with tools, 92.7% on AIME 2025, 86.0% on GPQA Diamond, and scores 45.3 on Z.ai's internal coding suite, alongside a #1 placement on Vending Bench 2 among open-source models. Z.ai distributes the weights openly under an MIT-style license and frames GLM 5.1 as a flagship open alternative to closed frontier coding models, well suited for agentic engineering pipelines, research workflows, and production code generation where sustained autonomy and tool use matter more than single-turn chat.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Vercel AI Gateway
Zhipu is beginning to lean into monetization after a prolonged investment phase, raising prices for access to its most advanced AI models by at least 8% alongsi
Vercel AI Gateway
The Chinese company said its new open-source model can continue to improve over hundreds of iterations, as AI vendors race to build tools that can handle longer software tasks.
Vercel AI Gateway
You can now access GLM 5.1 on Vercel's AI Gateway with no markup and no other provider accounts required.
Vercel AI Gateway
InferenceX's profile provides technical context tying GLM-5.1 to its GLM-5 base, noting GLM-5 scales to 744B total parameters with 40B active and 28.5T pre-training tokens under an MIT license, and that GLM-5.1 is a follow-up point release on the same architecture released April 7, 2026, also MIT-licensed. It cites the The profile confirms that both GLM-5 and GLM-5.1 are served through the Z.ai API with a 200K context window and 128K maximum output, and links the GLM-5.1 weight repository (created April 3, 2026) to Hugging Face. The excerpt makes no claims about Vercel AI Gateway hosting or pricing for GLM-5.1.
Vercel AI Gateway
The Z.AI release notes index logs a 2026-04-07 entry for GLM-5.1 stating the model is designed for long-horizon tasks and can work independently for up to 8 hours in a single run, covering planning, execution, iterative refinement, and final delivery. The release-notes excerpt credits GLM-5.1's gains to multi-turn SFT, The same release-notes page also lists subsequent versions—GLM-5.2 (2026-06-16) introducing a 1M lossless context, and GLM-5.3 / GLM-5.3-Flash (2026-08-26) adding native visual capabilities and stronger coding and cybersecurity results—confirming GLM-5.1's place as the April 2026 baseline within Z.AI's rapidly evolving