Sulat.com
AI models
Merge Gateway logo

Model details

GLM-4.5 Air

GLM-4.5 Air is the streamlined sibling in Zhipu AI's GLM-4.5 family, a Mixture-of-Experts architecture that activates roughly 12 billion of its 106 billion total parameters per inference pass. This sparsity is the foundation of its efficiency: rather than running every weight for every token, the model routes work through a small active subset, enabling faster responses and lower compute costs while preserving much of the capacity of the flagship 355-billion-parameter GLM-4.5. The family's design philosophy emphasizes natively integrating reasoning, coding, and agentic abilities into a single unified model, making the Air variant well suited as a workhorse for tool-using and function-calling pipelines where developers want agent behavior without flagship-scale expense.

In practical terms, GLM-4.5 Air shines on agent evaluation benchmarks. On Galileo's Agent Leaderboard it posts a 0.940 Tool Selection Quality score, while costing roughly 94% less than Claude Sonnet 4.5 and delivering an average response latency of about 0.64 seconds. Galileo also reports a blended token rate around $0.42 per million tokens, positioning the model as a strong fit for high-volume agent deployments, function-calling workflows, and large-scale tool orchestration where budget efficiency matters more than pushing the absolute frontier of general knowledge. Developers should weigh that efficiency against its trailing performance on broad knowledge benchmarks, where it sits notably behind top proprietary peers, and against documented brittleness in specialized domains. Overall, the model is a pragmatic choice when agent reliability and cost per call are the binding constraints.

Merge Gatewayzai/glm-4.5-airglm-air

Quick Info

Powered by
Provider
Merge Gateway
Model key
zai/glm-4.5-air
Release date
Jul 28, 2025
Last updated
Jul 28, 2025
Knowledge cutoff
2025-04
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.20
Output token cost
$1.10

Limits

Output tokens
98,304 tokens
Context window
128,000 tokens

Transparent token rates

Compare glm-air pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GLM-4.5 Air

Videos about GLM-4.5 Air

More models around GLM-4.5 Air