Sulat.com
AI models
NanoGPT logo

Model details

GLM 4.7 Flash Thinking

The supplied evidence only supports the existence of GLM-4.7 Flash as a smaller sibling to the 384B-parameter GLM-4.7. An independent technical analysis describes it as a 30B-parameter mixture-of-experts design with roughly 3B parameters active for each generated token. That structure is intended to reduce the compute and memory demands of the larger model while preserving substantial capacity, making the base model relevant to complex language-generation and reasoning workloads.

The same analysis identifies multi-head latent attention and multi-token prediction as efficiency-oriented architectural features. Multi-head latent attention is designed to lower the memory needed for key-value caching, while multi-token prediction is intended to improve generation throughput; together, they make the model more practical for suitably equipped local or private deployments than its full-sized sibling. The write-up indicates that a 4-bit representation can make the model viable on high-memory workstation hardware, but it does not provide directly extractable benchmark results or verify a distinct thinking-specific variant.

NanoGPTz-ai/glm-4.7-flash:thinkingglm-flash

Quick Info

Powered by
Provider
NanoGPT
Model key
z-ai/glm-4.7-flash:thinking
Release date
Jan 19, 2026
Last updated
Jan 19, 2026
Knowledge cutoff
2025-04
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.07
Output token cost
$0.40

Limits

Input tokens
200,000 tokens
Output tokens
128,000 tokens
Context window
200,000 tokens

Latest news about GLM 4.7 Flash Thinking

Videos about GLM 4.7 Flash Thinking

Recent tweets and retweets from NanoGPT

More models around GLM 4.7 Flash Thinking