Sulat.com
AI models
Amazon Bedrock logo

Model details

GLM-4.7-Flash

GLM-4.7 Flash takes a practical dense-architecture approach rather than the massive MoE design of its flagship sibling. With its 31-billion-parameter dense structure, the model activates all parameters for every token generated, producing highly predictable inference behavior without the latency spikes that can plague larger sparse systems. This architectural choice makes the model far easier to quantize and deploy locally, and it enables the model to hit 80+ tokens per second on consumer hardware like MacBooks—a meaningful figure for developers who want meaningful AI assistance without renting time from massive proprietary systems. The design philosophy here is pragmatic workhorse rather than all-knowing AGI: the model is explicitly built for teams and individuals who need reliable code generation and reasoning but are constrained by budget or hardware.

The model shows its strongest identity in coding and agentic scenarios, achieving 59% on Software Engineering Bench, 79.5% on TA2 agentic tasks, and 75.2% on GPQA—scores that position it as a credible alternative where proprietary models would otherwise dominate budgets. Released under an MIT license with free API access, GLM-4.7 Flash is designed for the full development lifecycle: from experimenting with ideas to shipping real tools, making agentic workflows viable for smaller teams that could never justify the cost of massive frontier models. The 128,000-token context window allows developers to feed entire repositories, heavy documentation, or lengthy stack traces into a single conversation, and the dense architecture simplifies local quantization for teams that want to run their own infrastructure. This combination of open weights, coding-focused benchmarks, and hardware-friendly efficiency makes the model a natural fit for developers who want frontier-level coding ability without frontier-level costs.

Amazon Bedrockzai.glm-4.7-flashglm-flash

Quick Info

Powered by
Provider
Amazon Bedrock
Model key
zai.glm-4.7-flash
Release date
Jan 19, 2026
Last updated
Jan 19, 2026
Knowledge cutoff
2025-04
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.07
Output token cost
$0.40

Limits

Output tokens
131,072 tokens
Context window
200,000 tokens

Transparent token rates

Compare glm-flash pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GLM-4.7-Flash

Amazon Bedrock

CoverageBenchmark

Z.AI's GLM-4.7 Flash is a 31-billion-parameter open-source model released for coding, reasoning, and agentic workflows, offering free API access and local deployment. It records 59% on Software Engineering Bench, 79.5% on TA2 agentic tasks, and 75.2% on GPQA, while pricing lists $0.07 input, $0.01 cached input, and $0.

Amazon Bedrock

CoverageBenchmark

GLM-4.7 Flash packs 31B parameters and an MIT license with free API access, helping you test ideas and ship tools on a tiny budget.

Amazon Bedrock

Coverage

The 30B model achieves 59.2% on SWE-bench Verified while running at 80+ tokens per second on MacBooks.

Amazon Bedrock

Coverage

The news blog specialized in Japanese culture, odd news, gadgets and all other funny stuffs. Updated everyday.

Amazon Bedrock

CoverageRelease Notes

GLM-4.7-Flash is a new member of the GLM 4.7 family and targets developers who want strong coding and reasoning performance in a model that is practical to...

Videos about GLM-4.7-Flash

More models around GLM-4.7-Flash