Sulat.com
AI models
Cloudflare Workers AI logo

Model details

GLM-4.7-Flash

GLM-4.7-Flash is a 30-billion-parameter mixture-of-experts model with 3 billion active parameters per pass, designed as the most capable option in the 30B class while keeping inference economical. Zhipu AI released it as an open-weight model on Hugging Face, accompanied by a technical blog and the GLM-4.5 technical report, so developers can study the architecture lineage and run it locally or through hosted APIs. The Cloudflare Workers AI listing frames it as a fast multilingual text generator with a 131,072-token context window, optimized for dialogue, instruction following, and multi-turn tool calling across more than one hundred languages, which makes it a practical fit for global conversational agents and workflow automation.

The model's published benchmark profile points to a balanced emphasis on reasoning, coding, and agentic behavior rather than a narrow specialty. Reported scores from Zhipu's evaluation include 91.6 on AIME 25, 75.2 on GPQA, 64.0 on LiveCodeBench v6, 59.2 on SWE-bench Verified, 79.5 on τ²-Bench, and 42.8 on BrowseComp, measured against Qwen3-30B-A3B-Thinking-2507 and GPT-OSS-20B as reference peers. A standout result is the τ²-Bench agentic score, which is paired with a "Preserved Thinking" mode requirement for multi-turn tasks like τ²-Bench and Terminal Bench 2, indicating that sustained chain-of-thought context is treated as a first-class capability for tool-using workflows. Taken together, these traits make GLM-4.7-Flash a strong match for lightweight deployments that need open-weight flexibility, long-context reasoning, and reliable multi-step tool orchestration.

Cloudflare Workers AI@cf/zai-org/glm-4.7-flashglm-flash

Quick Info

Powered by
Provider
Cloudflare Workers AI
Model key
@cf/zai-org/glm-4.7-flash
Release date
Jan 19, 2026
Last updated
Jan 19, 2026
Knowledge cutoff
2025-04
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.0605
Output token cost
$0.40

Limits

Output tokens
131,072 tokens
Context window
131,072 tokens

Latest news about GLM-4.7-Flash

Videos about GLM-4.7-Flash

More models around GLM-4.7-Flash