Currently listed through these providers:
Model details
GLM-4.7-Flash
GLM-4.7 Flash is a dense, open-weight language model designed as a pragmatic alternative to massive proprietary coding systems. Unlike the flagship GLM-4.7, which relies on a 355-billion-parameter Mixture-of-Experts architecture, Flash uses a streamlined 30-billion-parameter dense structure that engages all parameters for every token processed. This design choice produces highly predictable inference behavior with no erratic latency spikes, making it considerably easier to deploy and balance on constrained hardware. The model targets teams that need strong code generation without the operational headaches or expense of renting frontier-scale systems.
Because Flash activates its full parameter set on every forward pass, it behaves more like a traditional dense transformer than its MoE sibling, which simplifies caching, scheduling, and capacity planning for production workloads. The variant targets efficient agentic coding workflows where predictable latency and stable throughput matter more than raw parameter count. An NVFP4-quantized build was released alongside the base weights, aimed at newer inference stacks such as Transformers 5.0 and vLLM 0.14, signaling continued investment in optimized deployment formats for teams running on modern GPU hardware.
Quick Info
Powered by- Provider
- Pendra
- Model key
- glm-4.7-flash
- Release date
- Jan 19, 2026
- Last updated
- Jan 19, 2026
- Knowledge cutoff
- 2025-04
- Input modalities
- Output modalities
- Capabilities
Cost
A provider subscription or plan supersedes token-based pricing for this model.
Limits
- Output tokens
- 131,072 tokens
- Context window
- 200,000 tokens
Latest news about GLM-4.7-Flash
Videos about GLM-4.7-Flash
More models around GLM-4.7-Flash
This exact model name is also listed by 17 other providers.