Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
ZenMux logo

Model details

GLM 4.5 Air

GLM 4.5 Air is a lighter member of Z.ai's GLM-4.5 family, engineered specifically for intelligent-agent applications where reasoning, coding, and tool use need to coexist. It uses a Mixture-of-Experts architecture with 106 billion total parameters and 12 billion active parameters, a more compact configuration than the flagship 355B/32B GLM-4.5. The design reflects a deliberate tradeoff: keeping expert-driven quality on complex tasks while lowering the inference cost and latency that agent loops demand.

A defining feature is its hybrid reasoning design, which lets a single deployment switch between a thinking mode for multi-step reasoning and tool orchestration and a non-thinking mode for immediate conversational responses, with reasoning toggled by an enable flag. Across twelve industry-standard benchmarks the GLM-4.5 series averaged 63.2, while GLM 4.5 Air delivered a competitive 59.8 with the efficiency advantages of its smaller active footprint, making it well suited for production agent pipelines, developer assistants, and retrieval or function-calling workflows that benefit from scalable reasoning on demand.

ZenMuxz-ai/glm-4.5-airglm-air

Quick Info

Powered by
Provider
ZenMux
Model key
z-ai/glm-4.5-air
Release date
Jul 28, 2025
Last updated
Jul 28, 2025
Knowledge cutoff
2025-04
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.1165
Output token cost
$0.2911

Limits

Output tokens
96,000 tokens
Context window
128,000 tokens

Transparent token rates

Compare GLM 4.5 Air pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GLM 4.5 Air

No articles yet. Fetch the latest news to show it here.

Videos about GLM 4.5 Air

More models around GLM 4.5 Air