Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Hugging Face logo

Model details

GLM-4.7-Flash

GLM-4.7-Flash is a 30B-class Mixture-of-Experts language model with around 3B active parameters, designed to balance strong reasoning performance with the efficiency needed for lightweight local deployment. The publisher describes it as the strongest model in the 30B tier, positioning it for scenarios where teams want a capable assistant without paying the cost of a much larger dense model. It is released as open weights and is widely accessible: the artifact is published under the zai-org organization on Hugging Face, mirrored in the Ollama library (where it has already attracted substantial downloads), and is also served as a managed API on the Z.ai platform, with usage guidance in the GLM-4.7 technical blog and reference to the GLM-4.5 technical report.

In benchmark terms, the model card reports results that are competitive with or ahead of comparably sized peers such as Qwen3-30B-A3B-Thinking-2507 and GPT-OSS-20B. Notable strengths include agentic and search-heavy evaluations, where it posts 59.2 on SWE-bench Verified, 79.5 on τ²-Bench, and 42.8 on BrowseComp, alongside solid reasoning scores of 91.6 on AIME 25, 75.2 on GPQA, and 64.0 on LCB v6. The publisher recommends enabling a Preserved Thinking mode for multi-turn agentic workloads like τ²-Bench and Terminal Bench 2, making the model a practical fit for developer tools, CLI coding agents, and retrieval or browser-driven assistants that need reasoning plus long-context handling rather than a maximally large general-purpose chat model.

Hugging Facezai-org/GLM-4.7-Flashglm-flash

Quick Info

Powered by
Provider
Hugging Face
Model key
zai-org/GLM-4.7-Flash
Release date
Aug 8, 2025
Last updated
Aug 8, 2025
Knowledge cutoff
2025-04
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
128,000 tokens
Context window
200,000 tokens

Latest news about GLM-4.7-Flash

Hugging Face

CoverageBenchmark

Zhipu AI (Z.ai) officially open-sourced GLM-4.7-Flash, a "Hybrid Thinking" model positioned as the strongest performer in the 30B class. It uses a 30B-A3B Mixture-of-Experts architecture, activating roughly 3B parameters per task to balance resource usage with processing power. Across key benchmarks, GLM-4.7-Flash reac The release emphasizes developer-friendly local deployment, with vLLM and SGLang already supporting the model on main (configurable via tensor-parallel-size, speculative-config, and the EAGLE algorithm) and Hugging Face transformers enabling direct invocation. The model targets agent applications in local or private-cl

Videos about GLM-4.7-Flash

More models around GLM-4.7-Flash