Currently listed through these providers:
Model details
GLM-4.5-Air
GLM-4.5-Air sits as the streamlined sibling within Z.ai's flagship GLM-4.5 family, purpose-built for agent-centric applications where responsiveness and reasoning depth both matter. It adopts the same Mixture-of-Experts architecture as the larger GLM-4.5 but with a markedly more compact footprint, carrying 106 billion total parameters of which 12 billion activate per inference. This design lets teams tap into frontier-style capabilities without provisioning for the full 355B-parameter variant, making it a practical middle ground for production agent workflows that need strong tool orchestration alongside conversational fluency.
A defining trait of GLM-4.5-Air is its hybrid inference design, which exposes two complementary modes through a reasoning-enabled toggle. The thinking mode unlocks advanced reasoning and tool use for multi-step problem solving, while the non-thinking mode delivers immediate responses suited to real-time interaction. Model weights are openly published on Hugging Face under the zai-org organization, so developers can self-host, fine-tune, or audit the system. Combined with its substantial context window and text-to-text interface, the model fits well into pipelines that alternate between quick-turn dialogue and deeper deliberative passes.
Quick Info
Powered by- Provider
- DevPass (LLM Gateway)
- Model key
- glm-4.5-air
- Release date
- Jul 28, 2025
- Last updated
- Jul 28, 2025
- Knowledge cutoff
- 2025-04
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.13
- Output token cost
- $0.85
Limits
- Output tokens
- 98,304 tokens
- Context window
- 131,000 tokens
Transparent token rates
Compare glm-air pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about GLM-4.5-Air
Videos about GLM-4.5-Air
More models around GLM-4.5-Air
This exact model name is also listed by 14 other providers.