Sulat.com
AI models
Synthetic logo

Model details

GLM-4.7-Flash

GLM-4.7-Flash is the free, lightweight entry in Z.AI's GLM-4.7 family, sitting alongside the standard GLM-4.7 and the paid FlashX tier while sharing the same 200K context and 128K maximum output budget. Third-party coverage describes it as a roughly 30B-parameter MoE model (reported as 30B-A3B or 31B) distributed under an MIT license, making it practical for developers who want to run capable code-and-reasoning inference on their own hardware rather than relying on a hosted endpoint. Its design intent is everyday developer use: streaming text output, function calling for tool integration, and a permissive license that supports both experimentation and local deployment on consumer machines.

Independent benchmarks and hands-on reports position GLM-4.7-Flash as a strong coding model for its class, with Techloy citing a 59.2% score on SWE-bench Verified and sustained throughput above 80 tokens per second on MacBooks, while Geeky Gadgets frames its pricing as dramatically undercutting premium alternatives. Z.AI's own documentation highlights multiple thinking modes, streaming output, function calling, and context caching as supported capabilities, which together make the model well suited to multi-step agent tasks and interactive coding assistants. Active forum threads on NVIDIA's developer community already show users requesting AWQ and NVFP4 quantized builds for DGX Spark hardware, signaling fast grassroots adoption among developers targeting local or prosumer inference setups.

Synthetichf:zai-org/GLM-4.7-Flashglm-flash

Quick Info

Powered by
Provider
Synthetic
Model key
hf:zai-org/GLM-4.7-Flash
Release date
Jan 19, 2026
Last updated
Jan 19, 2026
Knowledge cutoff
2025-04
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.10
Output token cost
$0.50

Limits

Output tokens
65,536 tokens
Context window
196,608 tokens

Transparent token rates

Compare glm-flash pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GLM-4.7-Flash

Synthetic

CoverageBenchmark

GLM-4.7 Flash packs 31B parameters and an MIT license with free API access, helping you test ideas and ship tools on a tiny budget.

Synthetic

Coverage

The news blog specialized in Japanese culture, odd news, gadgets and all other funny stuffs. Updated everyday.

Synthetic

Coverage

The 30B model achieves 59.2% on SWE-bench Verified while running at 80+ tokens per second on MacBooks.

Synthetic

CoverageRelease Notes

GLM-4.7-Flash is a new member of the GLM 4.7 family and targets developers who want strong coding and reasoning performance in a model that is practical to...

Synthetic

Coverage

Excerpt): Zhipu AI has open-sourced GLM-4.7-Flash, a 30B-parameter MoE model that activates only 3B during inference. Notably, it debuts the MLA architecture for efficiency and runs at 43 tokens/sec on an Apple M5 laptop.

Synthetic

CoverageAnalysis

From Weights & Biases - deep dive into Clawdbot, an personal AI assistant that learns and evolves, GLM 4.7 Flash, a bunch of new TTS models and Claude's new constitution!

Videos about GLM-4.7-Flash

More models around GLM-4.7-Flash