Currently listed through these providers:
Model details
GLM 4.5
GLM-4.5 is built on a self-developed Mixture of Experts architecture, designed to integrate reasoning, software development, and agentic decision-making into a single, cohesive system. By moving away from specialized models that focus on isolated tasks, this architecture aims to provide a balanced intelligence capable of handling the multifaceted requirements of modern agentic applications. The flagship model utilizes 355 billion total parameters with 32 billion active parameters, while the streamlined Air variant operates with 106 billion total parameters and 12 billion active parameters, ensuring high performance across a wide range of real-world scenarios.
The model series features a hybrid reasoning design that allows users to toggle between a thinking mode for complex, multi-step problem solving and a non-thinking mode for instant, snappy responses. This flexibility makes it particularly effective for high-volume agent deployments and tool orchestration pipelines where both depth and speed are required. By leveraging its efficient parameter utilization, the series offers a practical path for developers to implement sophisticated agentic workflows that remain cost-effective without sacrificing the cognitive depth necessary for challenging technical and analytical tasks.
Quick Info
Powered by- Provider
- ZenMux
- Model key
- z-ai/glm-4.5
- Release date
- Jul 28, 2025
- Last updated
- Jul 28, 2025
- Knowledge cutoff
- 2025-04
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.2911
- Output token cost
- $1.1645
Limits
- Output tokens
- 96,000 tokens
- Context window
- 128,000 tokens
Transparent token rates
Compare GLM 4.5 pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about GLM 4.5
No articles yet. Fetch the latest news to show it here.