Sulat.com
AI models
Pendra logo

Model details

GLM-4.7-Flash

GLM-4.7 Flash is a dense, open-weight language model designed as a pragmatic alternative to massive proprietary coding systems. Unlike the flagship GLM-4.7, which relies on a 355-billion-parameter Mixture-of-Experts architecture, Flash uses a streamlined 30-billion-parameter dense structure that engages all parameters for every token processed. This design choice produces highly predictable inference behavior with no erratic latency spikes, making it considerably easier to deploy and balance on constrained hardware. The model targets teams that need strong code generation without the operational headaches or expense of renting frontier-scale systems.

Because Flash activates its full parameter set on every forward pass, it behaves more like a traditional dense transformer than its MoE sibling, which simplifies caching, scheduling, and capacity planning for production workloads. The variant targets efficient agentic coding workflows where predictable latency and stable throughput matter more than raw parameter count. An NVFP4-quantized build was released alongside the base weights, aimed at newer inference stacks such as Transformers 5.0 and vLLM 0.14, signaling continued investment in optimized deployment formats for teams running on modern GPU hardware.

Pendraglm-4.7-flashglm-flash

Quick Info

Powered by
Provider
Pendra
Model key
glm-4.7-flash
Release date
Jan 19, 2026
Last updated
Jan 19, 2026
Knowledge cutoff
2025-04
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
131,072 tokens
Context window
200,000 tokens

Latest news about GLM-4.7-Flash

Videos about GLM-4.7-Flash

More models around GLM-4.7-Flash