Currently listed through these providers:
Model details
GLM-4.7 FlashX
GLM-4.7 FlashX sits inside the glm-flash family and is positioned as an efficiency-oriented variant for fast reasoning, coding, and agent-driven workflows. Open-weight availability makes it attractive to teams that need to inspect or self-host weights, while the long context window supports use cases like multi-document analysis, retrieval-augmented agents, and extended tool conversations. Its placement in the Flash tier reflects a design goal of keeping latency and cost low relative to larger GLM variants, so it is most useful as a workhorse model in high-volume pipelines rather than as a frontier research model.
In practical terms, the model is suited to cost-sensitive automations, background tasks, and production agents where structured outputs, tool calling, and temperature control matter more than advanced reasoning controls. Registries note that features such as reasoning effort, verbosity, thinking levels, computer use, and deep research are not supported, so teams building agentic systems should expect a leaner control surface than higher-end GLM models. Provider attribution differs across listings, with some sources crediting Zhipu AI and others Z.ai, so integration teams should confirm provenance through official channels before committing to a deployment path.
Quick Info
Powered by- Provider
- Merge Gateway
- Model key
- zai/glm-4.7-flashx
- Release date
- Jan 19, 2026
- Last updated
- Jan 19, 2026
- Knowledge cutoff
- 2025-04
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.07
- Output token cost
- $0.40
Limits
- Output tokens
- 131,072 tokens
- Context window
- 200,000 tokens
Transparent token rates
Compare glm-flash pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about GLM-4.7 FlashX
Videos about GLM-4.7 FlashX
More models around GLM-4.7 FlashX
This exact model name is also listed by 7 other providers.