Sulat.com
AI models
DevPass (LLM Gateway) logo

Model details

GLM-4.7-Flash

GLM-4.7-Flash is positioned as the efficiency-focused tier of the GLM-4.7 Series, marketed by Z.AI as a lightweight and free offering designed to balance performance with low-cost inference. Official documentation highlights upgrades in programming capability and more stable multi-step reasoning and execution, with a particular emphasis on handling complex agent tasks while still delivering natural conversational responses. The Ollama library listing reinforces this framing by describing the model as the strongest in the 30B class for lightweight deployment, suggesting a parameter footprint that supports local and edge-style serving without sacrificing core reasoning ability.

In practical terms, GLM-4.7-Flash is well suited to developer-facing workflows that need responsive inference on commodity hardware, including integrations with coding agents such as Claude Code, OpenCode, Hermes Agent, and OpenClaw that are demonstrated alongside the Ollama distribution. The model is openly available, with community-quantized GGUF builds redistributed through Hugging Face, making it attractive for self-hosted setups where data control or cost predictability matters. However, independent release tracking notes that this Flash variant has since been superseded by several newer Z.ai releases in the GLM-5.x line, so teams evaluating it today should weigh its maturity and ecosystem support against the latest alternatives for their specific latency and agent requirements.

DevPass (LLM Gateway)glm-4.7-flashglm-flash

Quick Info

Powered by
Provider
DevPass (LLM Gateway)
Model key
glm-4.7-flash
Release date
Jan 19, 2026
Last updated
Jan 19, 2026
Knowledge cutoff
2025-04
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.06
Output token cost
$0.40

Limits

Output tokens
131,072 tokens
Context window
200,000 tokens

Transparent token rates

Compare glm-flash pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GLM-4.7-Flash

DevPass (LLM Gateway)

CoverageRelease Notes

The Opper AI release tracker dates GLM-4.7-Flash to 19 January 2026 and lists it at $0.06 per million input tokens and $0.40 per million output tokens with a 203K context window—substantially undercutting the GLM-4.7 flagship and the subsequent GLM-5.x releases on per-token economics. An Artificial Analysis intelligenc More importantly, the tracker shows GLM-4.7-Flash has been superseded by six subsequent Z.ai releases: GLM-5 (Feb), GLM-5-Turbo (Mar), GLM-5.1/5V-Turbo (Apr), GLM-5.2 (Jun), GLM-5.3 (Aug 18), and GLM-5.3-Flash (Aug 26). The latest GLM-5.3-Flash ships with 1M multimodal context at $0.15/$0.50 per million tokens. Anyone

LLM Gateway

CoverageBenchmark

Thanks for sharing these unsloth quants!

Videos about GLM-4.7-Flash

More models around GLM-4.7-Flash