Currently listed through these providers:
Model details
GLM-5
GLM-5 is a next-generation large language model built for complex systems engineering and long-horizon agentic workflows. Compared to its predecessor, the model scales from 355 billion parameters with 32 billion active to 744 billion parameters with 40 billion active, while pre-training data expands from 23 trillion to 28.5 trillion tokens. To keep deployment practical at this scale, GLM-5 integrates DeepSeek Sparse Attention, which substantially reduces computational overhead while maintaining long-context capacity. The architecture is explicitly designed to move beyond generating code snippets or prototypes toward building complete systems and carrying out end-to-end engineering tasks.
The model advances post-training through slime, an asynchronous reinforcement learning infrastructure that improves training throughput and enables more granular iterations. This RL-driven approach bridges the gap between general competence and excellence, pushing the model toward frontier-level performance. On benchmarks across reasoning, coding, and agentic tasks, GLM-5 achieves best-in-class results among all open-source models available, closing the gap with leading proprietary systems. Released as open weights, it is available through Hugging Face, Amazon Bedrock, and NVIDIA NIM containers, making it accessible for both research and production deployment in autonomous coding and engineering workflows.
Quick Info
Powered by- Provider
- DInference
- Model key
- glm-5
- Release date
- Feb 12, 2026
- Last updated
- Feb 12, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.75
- Output token cost
- $2.40
Limits
- Output tokens
- 128,000 tokens
- Context window
- 200,000 tokens
Transparent token rates
Compare GLM-5 pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about GLM-5
No articles yet. Fetch the latest news to show it here.