GLM-4.5-Air is the streamlined sibling in the GLM-4.5 family, purpose-built as a foundation model for intelligent agent applications. It uses a Mixture-of-Experts architecture with 106 billion total parameters and 12 billion active per forward pass, trading raw scale for efficiency while keeping the same long context window as its larger counterpart. The model unifies reasoning, coding, and agent-oriented capabilities so a single system can plan, call tools, browse the web, and work through software engineering tasks. A defining design choice is its hybrid reasoning interface, which lets developers toggle between a Thinking Mode for multi-step problem solving and tool use, and a Non-Thinking Mode for fast, direct responses, with dynamic reasoning enabled by default.
GLM-4.5-Air follows the same training pipeline as the flagship GLM-4.5, starting from a large pretraining phase on 15 trillion tokens of general-domain data and then moving through targeted fine-tuning on code, reasoning, and agent-specific datasets, with reinforcement learning applied to sharpen those skills further. The weights, including FP8 versions, are released openly under the MIT license, and in a broad 12-benchmark evaluation the model lands at a score of 59.8, placing it in strong company with leading proprietary and open systems while running with notably less compute per token. In practice that combination makes GLM-4.5-Air a good fit for cost-conscious deployments that still need serious agent behavior, from front-end generation and software engineering assistants to retrieval-augmented chatbots and complex multi-step automation that can grow into more demanding workflows.