Currently listed through these providers:
Model details
GLM 4.7 Flash
GLM 4.7 Flash is a 30B-class mixture-of-experts model from the GLM family, designed as a lightweight option that balances performance and efficiency for local and hosted use. It is distributed with open weights through the zai-org organization on Hugging Face, accompanied by API access and a chat interface on the Z.ai platform. Default sampling parameters target temperature 1.0 and top-p 0.95 with a generous token budget, while a Preserved Thinking mode is recommended for multi-turn agentic evaluations such as τ²-Bench and Terminal Bench 2.
Benchmark results published on the model card position GLM 4.7 Flash competitively within its size class, reporting 91.6 on AIME 25, 75.2 on GPQA, 64.0 on LCB v6, 14.4 on HLE, 59.2 on SWE-bench Verified, 79.5 on τ²-Bench, and 42.8 on BrowseComp, compared against Qwen3-30B-A3B-Thinking-2507 and GPT-OSS-20B. The model is also packaged on the Ollama library, where it integrates with agent harnesses like Claude Code, OpenCode, and Hermes Agent through simple launch commands, making it a practical fit for developers building tool-using assistants and coding workflows that benefit from an efficient open-weights backbone.
Quick Info
Powered by- Provider
- Venice AI
- Model key
- zai-org-glm-4.7-flash
- Release date
- Jan 29, 2026
- Last updated
- Jun 11, 2026
- Knowledge cutoff
- 2025-04
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.06
- Output token cost
- $0.40
Limits
- Output tokens
- 16,384 tokens
- Context window
- 128,000 tokens
Transparent token rates
Compare GLM 4.7 Flash pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about GLM 4.7 Flash
No articles yet. Fetch the latest news to show it here.
Videos about GLM 4.7 Flash
More models around GLM 4.7 Flash
This exact model name is also listed by 8 other providers.