Currently listed through these providers:
Model details
GLM 4.5
GLM-4.5 was designed from the ground up as an agent-native foundation model, bringing together reasoning, coding, and tool-use capabilities in a unified architecture. Built on a Mixture of Experts framework that activates 32 billion of its 355 billion total parameters, it delivers efficient inference while maintaining deep computational power. The model introduces dual-mode operation—a thinking mode for multi-step reasoning and complex decision-making, and a non-thinking mode for rapid, real-time responses. This flexibility lets developers toggle between depth and speed depending on the task at hand, whether that's orchestrating autonomous agents, refactoring large codebases, or handling multi-turn dialogues with native tool calling.
The training pipeline scales up agentic capabilities through domain-specific and reasoning-focused pretraining, accumulating 30 trillion tokens across general, specialized, and code-focused data. Reinforcement learning further hones its reasoning, coding, and agent alignment—without requiring traditional feedback signals. Technical innovations include Grouped-Query Attention, Rotary Position Embedding for context handling, a Muon optimizer, and Multi-Token Prediction for faster generation. Performance benchmarks reflect this investment: 98.2% on MATH 500, 72.9% on LiveCodeBench, and strong tool selection quality scores on agent leaderboards. The combination of open-weight availability, MIT licensing, and a design philosophy centered on autonomous agent workloads positions GLM-4.5 as a compelling option for developers building self-improving agents, automated coding pipelines, and complex orchestration systems that need both power and adaptability.
Quick Info
Powered by- Provider
- Vercel AI Gateway
- Model key
- zai/glm-4.5
- Release date
- Jul 28, 2025
- Last updated
- Jul 28, 2025
- Knowledge cutoff
- 2025-07
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.60
- Output token cost
- $2.20
Limits
- Output tokens
- 96,000 tokens
- Context window
- 128,000 tokens
Transparent token rates
Compare GLM 4.5 pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about GLM 4.5
No articles yet. Fetch the latest news to show it here.