Currently listed through these providers:
Model details
GLM 4.7 Flash
GLM-4.7 Flash is a 30-billion-parameter Mixture-of-Experts model with about 3 billion active parameters, published by zai-org as the lightest member of the GLM-4.7 family. The Hugging Face model card positions it as the strongest model in the 30B class, aimed at lightweight deployment scenarios where the goal is balancing efficiency with strong reasoning performance. The card ties the release into a broader Z.ai ecosystem, linking to a dedicated GLM-4.7 technical blog, a GLM-4.5 technical report on arXiv, official API documentation on the Z.ai platform, and a chat.z.ai playground for quick experimentation.
On the published benchmark table, GLM-4.7 Flash posts competitive scores against two open peers in the same size band: AIME 25 at 91.6, GPQA at 75.2, HLE at 14.4, SWE-bench Verified at 59.2, τ²-Bench at 79.5, and BrowseComp at 42.8, with LCB v6 at 64.0. The card highlights especially large margins on agentic and search-heavy evaluations such as SWE-bench Verified, τ²-Bench, and BrowseComp, reflecting the model's focus on tool-using, multi-step coding workflows. For practical use, default decoding settings are temperature 1.0, top-p 0.95, and a 131,072-token generation ceiling, while multi-turn agentic tasks on τ²-Bench and Terminal Bench 2 are recommended to run with the "Preserved Thinking" mode to maintain coherent long-horizon behavior.
Quick Info
Powered by- Provider
- EmpirioLabs AI
- Model key
- glm-4-7-flash
- Release date
- Jan 19, 2026
- Last updated
- Jan 19, 2026
- Knowledge cutoff
- 2025-04
- Input modalities
- Output modalities
- Capabilities
Cost
A provider subscription or plan supersedes token-based pricing for this model.
Limits
- Output tokens
- 131,072 tokens
- Context window
- 200,000 tokens
Latest news about GLM 4.7 Flash
Videos about GLM 4.7 Flash
More models around GLM 4.7 Flash
This exact model name is also listed by 17 other providers.