Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Poe logo

Model details

glm-4.7-flash

GLM-4.7 Flash is a compact reasoning model built around a 30-billion parameter Mixture-of-Experts architecture, where roughly 3 billion parameters activate per token during inference. This design delivers strong capability in a much smaller footprint than the 355-billion parameter flagship, making it a practical choice for teams that need solid code generation without demanding massive hardware. The model leans into coding as a core strength, supporting multi-turn tool calling across more than a hundred languages, instruction-following, and dialogue tasks. Its architecture also offers predictable latency and simplifies quantization for teams deploying on their own infrastructure. While the context window is shorter than the flagship version, 128k tokens still handles full repository contexts, lengthy documentation, and complex stack traces comfortably for most real-world workflows.

The model emerged from Z.AI's GLM-4.7 series with benchmark results that place it among the strongest in its weight class, particularly on software engineering tasks and agentic evaluations. On SWE-bench Verified, which tests real-world bug resolution, it scores notably higher than comparable models in the 30B range, and on τ²-Bench, an agentic benchmark, it demonstrates strong multi-turn interaction capabilities that make it well-suited for tool-augmented workflows. Benchmarks like AIME 25 for mathematical reasoning and GPQA for domain knowledge also show competitive performance. The model ships as an open-weights release under an MIT license, with GGUF and MLX formats available for local deployment through tools like LM Studio, and requires around 16 GB of memory at minimum. This combination of strong coding performance, reasoning capability, and accessible deployment options positions it as a workhorse for teams building agentic pipelines or running code-heavy tasks at scale.

Poenovita/glm-4.7-flash

Quick Info

Powered by
Provider
Poe
Model key
novita/glm-4.7-flash
Release date
Jan 19, 2026
Last updated
Jan 19, 2026
Input modalities
Output modalities
Capabilities

Limits

Output tokens
65,500 tokens
Context window
200,000 tokens

Latest news about glm-4.7-flash

No articles yet. Fetch the latest news to show it here.

Videos about glm-4.7-flash