Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Zhipu AI logo

Model details

GLM-4.7-Flash

GLM-4.7-Flash is a 30B-class open-weight model positioned by its publisher as the strongest option in its parameter tier for lightweight deployment that balances performance and efficiency, and it is widely distributed through community runtimes like Ollama where it has already reached more than a million pulls. Third-party coverage describes it as a 31-billion-parameter, MIT-licensed release aimed at coding, reasoning, and agentic workflows, with free API access and the ability to run locally on consumer laptops, reportedly sustaining 80 or more tokens per second on MacBook hardware. That combination of permissive licensing, modest footprint, and agent-friendly tooling makes it attractive to developers who want to self-host coding assistants and experiment with agent pipelines without depending on a proprietary cloud service.

Benchmark reporting from independent outlets gives a clear picture of where GLM-4.7-Flash sits in the open-weight landscape. Coverage cites roughly 59 percent on SWE-bench Verified, 79.5 percent on TA2 agentic tasks, and 75.2 percent on GPQA, framing it as competitive with or ahead of similarly sized open peers such as gpt-oss-20b for programming-heavy workloads. Practical fit centers on local coding assistants, rapid prototyping of agent applications, and cost-sensitive API usage where its free or near-free input pricing and open weights keep iteration cheap, while longer-context tasks benefit from its documented context window without requiring frontier-class hardware.

Zhipu AIglm-4.7-flashglm-flash

Quick Info

Powered by
Provider
Zhipu AI
Model key
glm-4.7-flash
Release date
Jan 19, 2026
Last updated
Jan 19, 2026
Knowledge cutoff
2025-04
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
131,072 tokens
Context window
200,000 tokens

Latest news about GLM-4.7-Flash

Zhipu AI

CoverageBenchmark

GLM-4.7 Flash packs 31B parameters and an MIT license with free API access, helping you test ideas and ship tools on a tiny budget.

Zhipu AI

CoverageBenchmark

Z.AI's GLM-4.7 Flash is a 31-billion-parameter open-source model released for coding, reasoning, and agentic workflows, offering free API access and local deployment. It records 59% on Software Engineering Bench, 79.5% on TA2 agentic tasks, and 75.2% on GPQA, while pricing lists $0.07 input, $0.01 cached input, and $0.

Zhipu AI

Coverage

The 30B model achieves 59.2% on SWE-bench Verified while running at 80+ tokens per second on MacBooks.

Zhipu AI

Coverage

The news blog specialized in Japanese culture, odd news, gadgets and all other funny stuffs. Updated everyday.

Zhipu AI

Coverage

Excerpt): Zhipu AI has open-sourced GLM-4.7-Flash, a 30B-parameter MoE model that activates only 3B during inference. Notably, it debuts the MLA architecture for efficiency and runs at 43 tokens/sec on an Apple M5 laptop.

Zhipu AI

Coverage

Today,GLM-4.7-Flash has been officially released and open-sourced., Sign up to moomoo now! You deserve more than free commissions. Not only does moomoo offer free trading, we also offer you an intuitive mobile trading platform which will let you invest anywhere and anytime.

Zhipu AI

Coverage

A Chinese-language tech blog index on jikejun.com includes an entry (jikejun.com/355.html) titled "智谱AI掀桌子了!GLM-4.7-Flash正式开源:API永久免费,性能超越4.5" that claims Zhipu AI released and open-sourced GLM-4.7-Flash on January 20, 2026, with the framing of a "permanent free API" and performance exceeding GLM-4 The entry is the only supplied candidate that frames GLM-4.7-Flash specifically as an open-weight release from Zhipu AI with a free-API angle, which is a meaningfully different product-impact story than the pure pricing/benchmark comparison offered by llm-stats. However, because it is ~6 months stale relative to the Ju

Videos about GLM-4.7-Flash

More models around GLM-4.7-Flash