Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Vertex logo

Model details

GLM-5

GLM-5 represents a deliberate shift in AI development philosophy, designed from the ground up for complex systems engineering and long-horizon agentic tasks rather than simple code generation. The model scales to 744 billion total parameters with 40 billion active parameters during inference, integrating DeepSeek Sparse Attention to dramatically cut deployment costs while preserving the ability to handle extensive contexts. This architectural emphasis reflects a broader industry movement away from "vibe coding" toward what researchers call agentic engineering, where AI systems build complete, end-to-end software rather than isolated snippets or prototypes.

The model builds on its predecessor GLM-4.5 through both expanded pre-training, using 28.5 trillion tokens, and a novel post-training approach. Researchers developed "slime," an asynchronous reinforcement learning infrastructure that decouples generation from training to improve throughput and enable more granular alignment iterations. This investment in training efficiency appears to pay off: GLM-5 claims best-in-class performance among open-source models on reasoning, coding, and agentic benchmarks, and the technical report documents results that position the model alongside proprietary frontier systems in real-world software engineering challenges. Released under the MIT license, it offers developers a fully accessible path to deploy capable agentic AI without vendor lock-in.

Vertexzai-org/glm-5-maasglmdeprecated

Quick Info

Powered by
Provider
Vertex
Model key
zai-org/glm-5-maas
Release date
Feb 12, 2026
Last updated
Feb 12, 2026
AI SDK package
@ai-sdk/openai-compatible
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$1.00
Output token cost
$3.20

Limits

Output tokens
128,000 tokens
Context window
200,000 tokens

Transparent token rates

Compare GLM-5 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GLM-5

Vertex

CoverageBenchmark

The InferenceX profile "GLM-5 / GLM-5.1 — Architecture, Evals & Inference Performance" consolidates Z.ai's first-party model-card and blog material for the base GLM-5 release, whose Hugging Face weight repository was created on 2026-02-11. It documents the headline scale jump from GLM-4.5 — 744B total parameters with 4 The same InferenceX page positions GLM-5 as Z.ai's flagship model targeting complex systems engineering and long-horizon agentic tasks, served through the Z.ai API with a 200K context window and 128K maximum output. It also gives substantive coverage of the follow-up GLM-5.1 — same architecture, MIT-licensed, repositor

Videos about GLM-5

More models around GLM-5