Currently listed through these providers:
Model details
GLM-4.5
GLM-4.5 is built as a foundation model family for intelligent agents, combining reasoning, coding, and agentic capabilities in a unified system. It achieves this through a mixture-of-experts architecture with 355 billion total parameters and 32 billion activated parameters per token, allowing selective expert routing that keeps inference efficient while leveraging massive model capacity. The architecture incorporates grouped-query attention, rotary position embeddings, and multi-token prediction to handle complex, multi-step tasks. Users can toggle between thinking mode for deep reasoning through tool use and complex decision chains, and non-thinking mode for rapid, direct responses—making it adaptable across different workload demands.
The training pipeline draws on 15 trillion tokens of general pretraining data supplemented by 8 trillion domain-specific tokens and 7 trillion code and reasoning tokens, followed by reinforcement learning alignment across reasoning, code generation, and agent workflows. On industry benchmarks, GLM-4.5 scores 63.2, placing third among both proprietary and open-source models worldwide, while maintaining superior parameter efficiency compared to dense models of similar capability. Released under the MIT open-source license with both base and hybrid reasoning variants available, it supports up to 128K context and 96K output tokens per request. This combination of open access, dual-mode reasoning, and native tool calling positions GLM-4.5 well for high-volume agent deployments, autonomous code refactoring, and orchestrating complex multi-step workflows at scale.
Quick Info
Powered by- Provider
- 302.AI
- Model key
- glm-4.5
- Release date
- Jul 29, 2025
- Last updated
- Jul 29, 2025
- Knowledge cutoff
- 2025-04
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.286
- Output token cost
- $1.142
Limits
- Output tokens
- 98,304 tokens
- Context window
- 131,072 tokens
Transparent token rates
Compare GLM-4.5 pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about GLM-4.5
No articles yet. Fetch the latest news to show it here.