OpenRouter
Compare GLM 4.5 Air from Z.ai to other AI models on key metrics including benchmarks, price, context length, and other model features.
Model details
GLM-4.5-Air is built as a foundational model for agent-centric applications, leveraging a Mixture-of-Experts architecture to balance capability with efficiency. The model delivers 106 billion total parameters with 12 billion active parameters per forward pass, allowing it to handle reasoning, coding, and tool invocation tasks without the full computational cost of activating all weights. Its hybrid reasoning design offers two execution modes: a thinking mode for complex multi-step reasoning and tool use, and a non-thinking mode for real-time interactions. This flexibility makes it adaptable to diverse agent scenarios, from code-centric integrations in tools like Claude Code to arbitrary agent applications that rely on dynamic tool invocation and web browsing capabilities.
The model undergoes a two-stage training pipeline: pretraining on 15 trillion tokens of general-domain data followed by targeted fine-tuning on datasets covering code, reasoning, and agent-specific tasks. Reinforcement learning is applied to further enhance performance in complex reasoning, coding, and agent behaviors. The model achieves a benchmark score of 59.8, positioning it competitively among both proprietary and open-source alternatives while maintaining superior efficiency through its sparse activation design. Released under the MIT license with full open-weight access, GLM-4.5-Air serves developers building intelligent agents that require both strong foundational capabilities and the adaptability to handle real-world tool use and multi-step task completion.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
OpenRouter
Compare GLM 4.5 Air from Z.ai to other AI models on key metrics including benchmarks, price, context length, and other model features.
OpenRouter
GLM-4.5-Air is the lightweight variant of our latest flagship model family, also purpose-built for agent-centric applications. $0 per million input tokens, $0 per million output tokens. 131,072 token context window, maximum output of 96,000 tokens. Higher uptime with 4 providers. Includes independent benchmarks from Ar
This exact model name is also listed by 14 other providers.