Currently listed through these providers:
Model details
GLM 4.7 Flash Thinking
The supplied evidence only supports the existence of GLM-4.7 Flash as a smaller sibling to the 384B-parameter GLM-4.7. An independent technical analysis describes it as a 30B-parameter mixture-of-experts design with roughly 3B parameters active for each generated token. That structure is intended to reduce the compute and memory demands of the larger model while preserving substantial capacity, making the base model relevant to complex language-generation and reasoning workloads.
The same analysis identifies multi-head latent attention and multi-token prediction as efficiency-oriented architectural features. Multi-head latent attention is designed to lower the memory needed for key-value caching, while multi-token prediction is intended to improve generation throughput; together, they make the model more practical for suitably equipped local or private deployments than its full-sized sibling. The write-up indicates that a 4-bit representation can make the model viable on high-memory workstation hardware, but it does not provide directly extractable benchmark results or verify a distinct thinking-specific variant.
Quick Info
Powered by- Provider
- NanoGPT
- Model key
- z-ai/glm-4.7-flash:thinking
- Release date
- Jan 19, 2026
- Last updated
- Jan 19, 2026
- Knowledge cutoff
- 2025-04
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.07
- Output token cost
- $0.40
Limits
- Input tokens
- 200,000 tokens
- Output tokens
- 128,000 tokens
- Context window
- 200,000 tokens