Sulat.com
AI models
Z.AI logo

Model details

GLM-4.7-Flash

GLM-4.7-Flash is designed as a pragmatic, high-performance alternative to larger, more complex systems. While its flagship counterpart utilizes a massive Mixture-of-Experts architecture, this model employs a streamlined, dense 30-billion parameter structure. This design choice ensures that every parameter is utilized for every token processed, resulting in highly predictable inference behavior and consistent latency. By focusing on a dense architecture, the model offers a reliable and efficient solution for teams that require robust code generation and reasoning capabilities without the hardware management challenges often associated with larger, sparse models.

Built to serve as a workhorse for developers, the model excels in real-world coding tasks and complex tool-calling scenarios. Its architecture simplifies quantization and deployment on local hardware, making it a versatile choice for engineers looking to integrate agentic workflows into smaller-scale environments. With a substantial context window, it is well-suited for processing full repository contexts, extensive documentation, and detailed stack traces. As a dense model, it provides a stable foundation for those seeking a balance between high-level performance and the practical constraints of budget and infrastructure.

Z.AIglm-4.7-flashglm-flash

Quick Info

Powered by
Provider
Z.AI
Model key
glm-4.7-flash
Release date
Jan 19, 2026
Last updated
Jan 19, 2026
Knowledge cutoff
2025-04
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
131,072 tokens
Context window
200,000 tokens

Latest news about GLM-4.7-Flash

Z.AI

CoverageBenchmark

GLM-4.7 Flash packs 31B parameters and an MIT license with free API access, helping you test ideas and ship tools on a tiny budget.

Z.AI

CoverageBenchmark

Z.AI's GLM-4.7 Flash is a 31-billion-parameter open-source model released for coding, reasoning, and agentic workflows, offering free API access and local deployment. It records 59% on Software Engineering Bench, 79.5% on TA2 agentic tasks, and 75.2% on GPQA, while pricing lists $0.07 input, $0.01 cached input, and $0.

Z.AI

Coverage

The 30B model achieves 59.2% on SWE-bench Verified while running at 80+ tokens per second on MacBooks.

Z.AI

CoverageRelease Notes

Zhipu AI Releases GLM-4.7-Flash: A 30B-A3B MoE (Mixture of Experts) Model for Efficient Local Coding and Agents.

Z.AI

Coverage

Excerpt): Zhipu AI has open-sourced GLM-4.7-Flash, a 30B-parameter MoE model that activates only 3B during inference. Notably, it debuts the MLA architecture for efficiency and runs at 43 tokens/sec on an Apple M5 laptop.

Videos about GLM-4.7-Flash

More models around GLM-4.7-Flash