Currently listed through these providers:
Model details
GLM 4.7 Flash
GLM 4.7 Flash is positioned as a pragmatic alternative to massive proprietary coding models, designed by Z.AI as a 30-billion parameter dense architecture that prioritizes efficiency over scale. Unlike its larger sibling, which relies on a 355-billion parameter Mixture-of-Experts design, Flash uses a streamlined dense structure where every parameter is engaged on every token. This architectural choice translates into highly predictable inference behavior, avoiding the latency spikes and hardware headaches common with MoE systems, and makes the model considerably easier to deploy on constrained infrastructure.
The model's intended audience is teams that need reliable code generation without the operational complexity or cost of frontier-scale systems, making agentic coding workflows more accessible to smaller engineering groups. A community-quantized variant called GLM-4.7-Flash-NVFP4 has also surfaced on NVIDIA's DGX Spark / GB10 developer forum, targeting compatibility with Transformers 5.0 and vLLM 0.14, which suggests growing ecosystem momentum around efficient serving of this architecture. Combined with its dense design and coding-focused positioning, GLM 4.7 Flash fits naturally as a workhorse for DevOps and developer-tooling scenarios where predictable latency and straightforward deployment matter more than raw breadth.
Quick Info
Powered by- Provider
- NanoGPT
- Model key
- z-ai/glm-4.7-flash
- Release date
- Jan 19, 2026
- Last updated
- Jan 19, 2026
- Knowledge cutoff
- 2025-04
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.07
- Output token cost
- $0.40
Limits
- Input tokens
- 200,000 tokens
- Output tokens
- 128,000 tokens
- Context window
- 200,000 tokens
Latest news about GLM 4.7 Flash
Videos about GLM 4.7 Flash
Recent tweets and retweets from NanoGPT
More models around GLM 4.7 Flash
This exact model name is also listed by 17 other providers.