Currently listed through these providers:
Model details
GLM-4.5 Air
GLM-4.5 Air is the streamlined sibling in Zhipu AI's GLM-4.5 family, a Mixture-of-Experts architecture that activates roughly 12 billion of its 106 billion total parameters per inference pass. This sparsity is the foundation of its efficiency: rather than running every weight for every token, the model routes work through a small active subset, enabling faster responses and lower compute costs while preserving much of the capacity of the flagship 355-billion-parameter GLM-4.5. The family's design philosophy emphasizes natively integrating reasoning, coding, and agentic abilities into a single unified model, making the Air variant well suited as a workhorse for tool-using and function-calling pipelines where developers want agent behavior without flagship-scale expense.
In practical terms, GLM-4.5 Air shines on agent evaluation benchmarks. On Galileo's Agent Leaderboard it posts a 0.940 Tool Selection Quality score, while costing roughly 94% less than Claude Sonnet 4.5 and delivering an average response latency of about 0.64 seconds. Galileo also reports a blended token rate around $0.42 per million tokens, positioning the model as a strong fit for high-volume agent deployments, function-calling workflows, and large-scale tool orchestration where budget efficiency matters more than pushing the absolute frontier of general knowledge. Developers should weigh that efficiency against its trailing performance on broad knowledge benchmarks, where it sits notably behind top proprietary peers, and against documented brittleness in specialized domains. Overall, the model is a pragmatic choice when agent reliability and cost per call are the binding constraints.
Quick Info
Powered by- Provider
- Merge Gateway
- Model key
- zai/glm-4.5-air
- Release date
- Jul 28, 2025
- Last updated
- Jul 28, 2025
- Knowledge cutoff
- 2025-04
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.20
- Output token cost
- $1.10
Limits
- Output tokens
- 98,304 tokens
- Context window
- 128,000 tokens
Transparent token rates
Compare glm-air pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about GLM-4.5 Air
Videos about GLM-4.5 Air
More models around GLM-4.5 Air
This exact model name is also listed by 14 other providers.