Currently listed through these providers:
Model details
GLM-4.7-Flash
GLM-4.7-Flash is a 30B-A3B Mixture-of-Experts model trained by Z.ai, designed as the smaller, efficiency-minded sibling within the GLM-4.7 family. Rather than being a dense model, it relies on sparse expert activation, which lets it deliver stronger performance than a typical 30B-parameter design while keeping inference latency and compute costs manageable. The family itself is built on a new base model and is explicitly oriented toward coding and tool calling, with GLM-4.7-Flash positioned for lightweight deployment where teams want a balance between capability and resource use.
In practical terms, GLM-4.7-Flash targets developers who need an open-weight model that can reason step-by-step and integrate with external tools, rather than a general-purpose chat model. It supports tool use and reasoning out of the box, and its weights are distributed in gguf and mlx formats so they can run locally through tools like LM Studio, with around 16 GB of RAM cited as the minimum for the smallest GLM-4.7 variant. That combination of an MoE backbone, a coding- and agent-leaning design, and open distribution makes it a natural fit for self-hosted coding assistants, prototyping agent workflows, and other developer-facing scenarios where staying inside an open ecosystem matters more than squeezing out the last few points of benchmark performance.
Quick Info
Powered by- Provider
- Eden AI
- Model key
- deepinfra/zai-org/GLM-4.7-Flash
- Release date
- Jan 19, 2026
- Last updated
- Jan 19, 2026
- Knowledge cutoff
- 2025-04
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.06
- Output token cost
- $0.40
Limits
- Output tokens
- 131,072 tokens
- Context window
- 202,752 tokens
Latest news about GLM-4.7-Flash
Videos about GLM-4.7-Flash
More models around GLM-4.7-Flash
This exact model name is also listed by 17 other providers.