Currently listed through these providers:
Model details
GLM 4.7 FlashX
GLM-4.7 FlashX extends Zhipu AI's Mixture-of-Experts foundation, building on a 30-billion parameter architecture where only 3 billion parameters activate per token during inference. This sparse activation design allows the model to deliver enhanced performance without the memory and compute demands of a dense 30B model. The architecture positions the model as a high-performance variant within the GLM-4.7 Flash family, optimized for scenarios that require both capability and efficiency.
The model targets practical deployment across diverse hardware environments, with benchmarked speeds ranging from 43 to 81 tokens per second on consumer-grade hardware like RTX 3090s and Apple Silicon machines. Beyond straightforward text generation, the model has been designed for creative writing, translation, long-context comprehension, and role-play applications where extended context windows provide meaningful advantage. The FlashX designation indicates it represents an enhanced tier within the GLM lineup, offering faster performance characteristics compared to the base Flash variant for users requiring higher throughput while maintaining the lightweight efficiency the series is known for.
Quick Info
Powered by- Provider
- ZenMux
- Model key
- z-ai/glm-4.7-flashx
- Release date
- Jan 19, 2026
- Last updated
- Jan 19, 2026
- Knowledge cutoff
- 2025-04
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.0728
- Output token cost
- $0.4367
Limits
- Output tokens
- 128,000 tokens
- Context window
- 200,000 tokens
Transparent token rates
Compare GLM 4.7 FlashX pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about GLM 4.7 FlashX
No articles yet. Fetch the latest news to show it here.
Videos about GLM 4.7 FlashX
More models around GLM 4.7 FlashX
This exact model name is also listed by 3 other providers.