Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Venice AI logo

Model details

GLM 4.7 Flash

GLM 4.7 Flash is a 30B-class mixture-of-experts model from the GLM family, designed as a lightweight option that balances performance and efficiency for local and hosted use. It is distributed with open weights through the zai-org organization on Hugging Face, accompanied by API access and a chat interface on the Z.ai platform. Default sampling parameters target temperature 1.0 and top-p 0.95 with a generous token budget, while a Preserved Thinking mode is recommended for multi-turn agentic evaluations such as τ²-Bench and Terminal Bench 2.

Benchmark results published on the model card position GLM 4.7 Flash competitively within its size class, reporting 91.6 on AIME 25, 75.2 on GPQA, 64.0 on LCB v6, 14.4 on HLE, 59.2 on SWE-bench Verified, 79.5 on τ²-Bench, and 42.8 on BrowseComp, compared against Qwen3-30B-A3B-Thinking-2507 and GPT-OSS-20B. The model is also packaged on the Ollama library, where it integrates with agent harnesses like Claude Code, OpenCode, and Hermes Agent through simple launch commands, making it a practical fit for developers building tool-using assistants and coding workflows that benefit from an efficient open-weights backbone.

Venice AIzai-org-glm-4.7-flashglm-flash

Quick Info

Powered by
Provider
Venice AI
Model key
zai-org-glm-4.7-flash
Release date
Jan 29, 2026
Last updated
Jun 11, 2026
Knowledge cutoff
2025-04
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.06
Output token cost
$0.40

Limits

Output tokens
16,384 tokens
Context window
128,000 tokens

Transparent token rates

Compare GLM 4.7 Flash pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GLM 4.7 Flash

No articles yet. Fetch the latest news to show it here.

Videos about GLM 4.7 Flash

More models around GLM 4.7 Flash