GLM-4.7-Flash is a 30B-class Mixture-of-Experts language model with around 3B active parameters, designed to balance strong reasoning performance with the efficiency needed for lightweight local deployment. The publisher describes it as the strongest model in the 30B tier, positioning it for scenarios where teams want a capable assistant without paying the cost of a much larger dense model. It is released as open weights and is widely accessible: the artifact is published under the zai-org organization on Hugging Face, mirrored in the Ollama library (where it has already attracted substantial downloads), and is also served as a managed API on the Z.ai platform, with usage guidance in the GLM-4.7 technical blog and reference to the GLM-4.5 technical report.
In benchmark terms, the model card reports results that are competitive with or ahead of comparably sized peers such as Qwen3-30B-A3B-Thinking-2507 and GPT-OSS-20B. Notable strengths include agentic and search-heavy evaluations, where it posts 59.2 on SWE-bench Verified, 79.5 on τ²-Bench, and 42.8 on BrowseComp, alongside solid reasoning scores of 91.6 on AIME 25, 75.2 on GPQA, and 64.0 on LCB v6. The publisher recommends enabling a Preserved Thinking mode for multi-turn agentic workloads like τ²-Bench and Terminal Bench 2, making the model a practical fit for developer tools, CLI coding agents, and retrieval or browser-driven assistants that need reasoning plus long-context handling rather than a maximally large general-purpose chat model.