Sulat.com
AI models
NovitaAI logo

Model details

GLM 4.5 Air

GLM-4.5-Air is Z.AI's compact counterpart in the GLM-4.5 family, built from the ground up as a foundational model for agent-oriented applications. It shares its training lineage with the larger flagship: a broad pre-training phase followed by targeted fine-tuning on code, reasoning, and agent-specific datasets, with reinforcement learning layered on to sharpen those capabilities. The result is a model aimed squarely at workflows where the system has to plan, call tools, browse, write code, and chain together multi-step tasks rather than just answer isolated questions.

Architecturally, GLM-4.5-Air adopts a Mixture-of-Experts design with 106B total parameters and 12B active parameters per forward pass, trading some raw capacity for noticeably leaner inference. It runs in two modes, a thinking mode for complex reasoning and tool use, and a non-thinking mode for faster, more conversational responses, so developers can dial effort up or down depending on the task. Optimizations target tool invocation, web browsing, software engineering, and front-end development, making it a natural fit for code-centric agents like Claude Code and Roo Code as well as custom agent pipelines built through tool-calling APIs. Weights are openly published on Hugging Face under the zai-org namespace, which simplifies self-hosting, fine-tuning, and integration into private stacks.

NovitaAIzai-org/glm-4.5-airglm-air

Quick Info

Powered by
Provider
NovitaAI
Model key
zai-org/glm-4.5-air
Release date
Oct 13, 2025
Last updated
Oct 13, 2025
Knowledge cutoff
2025-04
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.13
Output token cost
$0.85

Limits

Output tokens
98,304 tokens
Context window
131,072 tokens

Transparent token rates

Compare glm-air pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GLM 4.5 Air

NovitaAI

CoverageBenchmark

The OpenRouter model page for z-ai/glm-4.5-air directly confirms NovitaAI as a hosting provider for the GLM 4.5 Air model, with listed pricing of $0.13 per million input tokens and $0.85 per million output tokens, a 131K context window, and a knowledge cutoff of December 2024. The model is described as the lightweight Provider telemetry on the same page shows NovitaAI delivering the best observed latency among GLM 4.5 Air hosts at 0.67 seconds P50 and the highest throughput at 51 tokens per second, with 99.56% uptime; effective (post-cache) pricing through NovitaAI is $0.03748 per million input tokens and $0.849 per million output,

Videos about GLM 4.5 Air

More models around GLM 4.5 Air