Synthetic
GLM-4.7 Flash packs 31B parameters and an MIT license with free API access, helping you test ideas and ship tools on a tiny budget.
Model details
GLM-4.7-Flash is the free, lightweight entry in Z.AI's GLM-4.7 family, sitting alongside the standard GLM-4.7 and the paid FlashX tier while sharing the same 200K context and 128K maximum output budget. Third-party coverage describes it as a roughly 30B-parameter MoE model (reported as 30B-A3B or 31B) distributed under an MIT license, making it practical for developers who want to run capable code-and-reasoning inference on their own hardware rather than relying on a hosted endpoint. Its design intent is everyday developer use: streaming text output, function calling for tool integration, and a permissive license that supports both experimentation and local deployment on consumer machines.
Independent benchmarks and hands-on reports position GLM-4.7-Flash as a strong coding model for its class, with Techloy citing a 59.2% score on SWE-bench Verified and sustained throughput above 80 tokens per second on MacBooks, while Geeky Gadgets frames its pricing as dramatically undercutting premium alternatives. Z.AI's own documentation highlights multiple thinking modes, streaming output, function calling, and context caching as supported capabilities, which together make the model well suited to multi-step agent tasks and interactive coding assistants. Active forum threads on NVIDIA's developer community already show users requesting AWQ and NVFP4 quantized builds for DGX Spark hardware, signaling fast grassroots adoption among developers targeting local or prosumer inference setups.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Synthetic
GLM-4.7 Flash packs 31B parameters and an MIT license with free API access, helping you test ideas and ship tools on a tiny budget.
Synthetic
The news blog specialized in Japanese culture, odd news, gadgets and all other funny stuffs. Updated everyday.
Synthetic
The 30B model achieves 59.2% on SWE-bench Verified while running at 80+ tokens per second on MacBooks.
Synthetic
GLM-4.7-Flash is a new member of the GLM 4.7 family and targets developers who want strong coding and reasoning performance in a model that is practical to...
Synthetic
Excerpt): Zhipu AI has open-sourced GLM-4.7-Flash, a 30B-parameter MoE model that activates only 3B during inference. Notably, it debuts the MLA architecture for efficiency and runs at 43 tokens/sec on an Apple M5 laptop.
Synthetic
From Weights & Biases - deep dive into Clawdbot, an personal AI assistant that learns and evolves, GLM 4.7 Flash, a bunch of new TTS models and Claude's new constitution!
This exact model name is also listed by 17 other providers.