Z.AI
GLM-4.7 Flash packs 31B parameters and an MIT license with free API access, helping you test ideas and ship tools on a tiny budget.
Model details
GLM-4.7-Flash is designed as a pragmatic, high-performance alternative to larger, more complex systems. While its flagship counterpart utilizes a massive Mixture-of-Experts architecture, this model employs a streamlined, dense 30-billion parameter structure. This design choice ensures that every parameter is utilized for every token processed, resulting in highly predictable inference behavior and consistent latency. By focusing on a dense architecture, the model offers a reliable and efficient solution for teams that require robust code generation and reasoning capabilities without the hardware management challenges often associated with larger, sparse models.
Built to serve as a workhorse for developers, the model excels in real-world coding tasks and complex tool-calling scenarios. Its architecture simplifies quantization and deployment on local hardware, making it a versatile choice for engineers looking to integrate agentic workflows into smaller-scale environments. With a substantial context window, it is well-suited for processing full repository contexts, extensive documentation, and detailed stack traces. As a dense model, it provides a stable foundation for those seeking a balance between high-level performance and the practical constraints of budget and infrastructure.
A provider subscription or plan supersedes token-based pricing for this model.
Z.AI
GLM-4.7 Flash packs 31B parameters and an MIT license with free API access, helping you test ideas and ship tools on a tiny budget.
Z.AI
Z.AI's GLM-4.7 Flash is a 31-billion-parameter open-source model released for coding, reasoning, and agentic workflows, offering free API access and local deployment. It records 59% on Software Engineering Bench, 79.5% on TA2 agentic tasks, and 75.2% on GPQA, while pricing lists $0.07 input, $0.01 cached input, and $0.
Z.AI
The 30B model achieves 59.2% on SWE-bench Verified while running at 80+ tokens per second on MacBooks.
Z.AI
Zhipu AI Releases GLM-4.7-Flash: A 30B-A3B MoE (Mixture of Experts) Model for Efficient Local Coding and Agents.
Z.AI
Excerpt): Zhipu AI has open-sourced GLM-4.7-Flash, a 30B-parameter MoE model that activates only 3B during inference. Notably, it debuts the MLA architecture for efficiency and runs at 43 tokens/sec on an Apple M5 laptop.
This exact model name is also listed by 17 other providers.