Currently listed through these providers:
Model details
GLM 5
GLM-5 is built as a large-scale Mixture of Experts model targeting complex systems engineering and long-horizon agentic tasks. The architecture scales dramatically from its predecessor—from 355B total parameters with 32B active to 744B total parameters with 40B active—enabling the model to handle more sophisticated reasoning challenges. A key technical choice is the integration of DeepSeek Sparse Attention, which preserves long-context capability while substantially reducing deployment costs. The model is explicitly designed to go beyond first-pass performance, staying effective across extended agentic sessions rather than exhausting its capabilities quickly. This makes it well-suited for ambiguous, multi-step problems where sustained judgment and repeated iteration matter more than quick initial responses.
The training pipeline builds on a significantly larger pre-training corpus of 28.5T tokens compared to 23T in earlier versions. A notable infrastructure advancement is slime, a custom asynchronous reinforcement learning framework that addresses the efficiency challenges of scaling RL for large language models. This enables more fine-grained post-training iterations that bridge the gap between basic competence and excellence. The combination of scaled pre-training, improved RL infrastructure, and architectural optimizations positions GLM-5 as a best-in-class performer among open-source models on reasoning, coding, and agentic benchmarks—closing the gap with frontier proprietary models. The model is available as open weights, supporting both research exploration and practical deployment for teams building autonomous coding or complex engineering agents.
Quick Info
Powered by- Provider
- ZenMux
- Model key
- z-ai/glm-5
- Release date
- Feb 12, 2026
- Last updated
- Feb 12, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.58
- Output token cost
- $2.60
Limits
- Output tokens
- 128,000 tokens
- Context window
- 200,000 tokens
Transparent token rates
Compare GLM 5 pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about GLM 5
No articles yet. Fetch the latest news to show it here.