Sulat.com
AI models
Eden AI logo

Model details

GLM-4.7-Flash

GLM-4.7-Flash is a 30B-A3B Mixture-of-Experts model trained by Z.ai, designed as the smaller, efficiency-minded sibling within the GLM-4.7 family. Rather than being a dense model, it relies on sparse expert activation, which lets it deliver stronger performance than a typical 30B-parameter design while keeping inference latency and compute costs manageable. The family itself is built on a new base model and is explicitly oriented toward coding and tool calling, with GLM-4.7-Flash positioned for lightweight deployment where teams want a balance between capability and resource use.

In practical terms, GLM-4.7-Flash targets developers who need an open-weight model that can reason step-by-step and integrate with external tools, rather than a general-purpose chat model. It supports tool use and reasoning out of the box, and its weights are distributed in gguf and mlx formats so they can run locally through tools like LM Studio, with around 16 GB of RAM cited as the minimum for the smallest GLM-4.7 variant. That combination of an MoE backbone, a coding- and agent-leaning design, and open distribution makes it a natural fit for self-hosted coding assistants, prototyping agent workflows, and other developer-facing scenarios where staying inside an open ecosystem matters more than squeezing out the last few points of benchmark performance.

Eden AIdeepinfra/zai-org/GLM-4.7-Flashglm-flash

Quick Info

Powered by
Provider
Eden AI
Model key
deepinfra/zai-org/GLM-4.7-Flash
Release date
Jan 19, 2026
Last updated
Jan 19, 2026
Knowledge cutoff
2025-04
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.06
Output token cost
$0.40

Limits

Output tokens
131,072 tokens
Context window
202,752 tokens

Latest news about GLM-4.7-Flash

Videos about GLM-4.7-Flash

More models around GLM-4.7-Flash