Sulat.com
AI models
Cortecs logo

Model details

GLM-4.7-Flash

GLM-4.7-Flash is designed as a pragmatic, high-performance alternative to larger, more complex systems. While its flagship counterpart utilizes a massive Mixture-of-Experts architecture, this model employs a streamlined, dense 30-billion parameter structure. This design choice ensures that every parameter is utilized for every token processed, resulting in highly predictable inference behavior and consistent latency. It is specifically engineered to serve as a reliable workhorse for developers who need robust code generation and repository-level context handling without the hardware overhead associated with larger, proprietary models.

The model is built to support agentic workflows and real-world programming tasks, offering a balance between performance and deployment flexibility. Because of its dense architecture, it is well-suited for quantization, making it a practical choice for teams looking to run local AI coding assistants on consumer hardware. With a 128,000-token context window, it effectively manages large codebases, documentation, and complex stack traces, providing a cost-effective solution for developers who require frontier-level coding chops in resource-constrained environments.

Cortecsglm-4.7-flashglm-flash

Quick Info

Powered by
Provider
Cortecs
Model key
glm-4.7-flash
Release date
Jan 19, 2026
Last updated
Jan 19, 2026
Knowledge cutoff
2025-04
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.08
Output token cost
$0.478

Limits

Output tokens
203,000 tokens
Context window
203,000 tokens

Transparent token rates

Compare glm-flash pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GLM-4.7-Flash

Cortecs

Coverage

The 30B model achieves 59.2% on SWE-bench Verified while running at 80+ tokens per second on MacBooks.

Videos about GLM-4.7-Flash

More models around GLM-4.7-Flash