Currently listed through these providers:
Model details
GLM-5
GLM 5 is a large-scale open-weight model designed to handle complex systems engineering and extended multi-step agent workflows. The architecture scales up to 744 billion parameters with a 40-billion-parameter active configuration, incorporating DeepSeek Sparse Attention to keep long-context performance intact while cutting deployment costs. Built with FlashAttention-4 optimizations for faster inference, the model targets developers and engineering teams who need reliable, sustained reasoning across coding tasks, system design challenges, and autonomous agent pipelines.
The model's development traces a lineage from GLM-4.5, growing pre-training data from 23 trillion to 28.5 trillion tokens to push further into reasoning and coding benchmarks. The team introduced SLIME, an asynchronous reinforcement learning infrastructure that dramatically improves post-training throughput and enables more granular optimization cycles. This combination of scaled pre-training and efficient RL-driven refinement helped GLM 5 achieve best-in-class standing among open-source models on academic benchmarks, closing the gap with frontier models on coding, reasoning, and agentic tasks. It is particularly well suited to developers building autonomous coding agents, long-horizon planning systems, and complex engineering workflows.
Quick Info
Powered by- Provider
- Cortecs
- Model key
- glm-5
- Release date
- Feb 12, 2026
- Last updated
- Feb 12, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.988
- Output token cost
- $3.164
Limits
- Output tokens
- 202,752 tokens
- Context window
- 202,752 tokens
Transparent token rates
Compare GLM-5 pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about GLM-5
No articles yet. Fetch the latest news to show it here.