Currently listed through these providers:
Model details
GLM-5
GLM-5 is designed for complex systems engineering and long-horizon agentic workflows, moving beyond conversational assistance toward autonomous software construction. Z.ai describes it as a step from "vibe coding" to agentic engineering, with the model offered through the Z.ai chat interface, the Z.ai API, and the Z.ai Coding Plan, alongside open weights on HuggingFace and a companion GitHub repository for the research community. A linked technical report on arXiv (2602.15763) provides deeper documentation of the training approach and evaluation results.
The model represents a substantial scale-up from its predecessor, expanding from 355B parameters with 32B active in the prior generation to 744B parameters with 40B active, and growing the pre-training corpus from 23T to 28.5T tokens. To keep this larger model economical at inference, Z.ai integrated DeepSeek Sparse Attention, which trims deployment cost while preserving long-context capacity for extended reasoning and codebase-scale inputs. Reinforcement learning was scaled through a custom asynchronous infrastructure called slime, enabling more fine-grained post-training iterations and contributing to marked gains on academic benchmarks relative to the earlier generation.
Quick Info
Powered by- Provider
- GMI Cloud
- Model key
- zai-org/GLM-5-FP8
- Release date
- Feb 12, 2026
- Last updated
- Feb 12, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.60
- Output token cost
- $1.92
Limits
- Output tokens
- 131,072 tokens
- Context window
- 202,752 tokens
Transparent token rates
Compare GLM-5 pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about GLM-5
No articles yet. Fetch the latest news to show it here.
