Currently listed through these providers:
Model details
GLM-4.5
GLM-4.5 is a flagship large language model built on a Mixture-of-Experts architecture with 355 billion total parameters and 32 billion activated parameters per token, designed to unify reasoning, coding, and agentic capabilities in a single system. The model was developed to address the growing complexity of agentic applications that demand both deep analytical thinking and real-time responsiveness. Its hybrid design offers two distinct inference modes: a thinking mode for complex reasoning chains and multi-step tool use, and a non-thinking mode for instant responses to simpler queries. The architecture incorporates Grouped-Query Attention, Rotary Position Embeddings, and Multi-Token Prediction, supported by the Muon optimizer for training efficiency.
The model was trained on approximately 30 trillion tokens combining general pretraining, domain-specific data, and dedicated code and reasoning corpora. GLM-4.5 represents a significant push toward making frontier-level AI accessible as open weights under an MIT license, distributing model weights through HuggingFace and ModelScope alongside API access. Its benchmark performance shows strong results on mathematical reasoning and code generation tasks, positioning it as a practical choice for developers building autonomous agents, coding assistants, and complex problem-solving applications. The design philosophy prioritizes agent-centric use cases, with native tool calling and multi-turn dialogue capabilities built into the core model rather than layered on as post-processing.
Quick Info
Powered by- Provider
- Z.AI
- Model key
- glm-4.5
- Release date
- Jul 28, 2025
- Last updated
- Jul 28, 2025
- Knowledge cutoff
- 2025-04
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.60
- Output token cost
- $2.20
Limits
- Output tokens
- 98,304 tokens
- Context window
- 131,072 tokens
Transparent token rates
Compare glm pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about GLM-4.5
No articles yet. Fetch the latest news to show it here.