Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Tempr Gateway logo

Model details

GLM-4.7-FlashX

The model overview is being prepared.

Tempr Gatewayzai/glm-4.7-flashxglm-flash

Quick Info

Powered by
Provider
Tempr Gateway
Model key
zai/glm-4.7-flashx
Release date
Jan 19, 2026
Last updated
Jan 19, 2026
Knowledge cutoff
2025-04
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.07
Output token cost
$0.40

Limits

Output tokens
131,072 tokens
Context window
200,000 tokens

Transparent token rates

Compare GLM-4.7-FlashX pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GLM-4.7-FlashX

Z.AI

Coverage

Reporting on the January 19, 2026 launch of the GLM-4.7 family, Techloy confirms that Z.AI offers GLM-4.7-FlashX as a faster paid companion to the free GLM-4.7-Flash, complementing Z.AI's local and API deployment options for coding workloads. The article positions FlashX within a tiered access strategy for developers. The same report situates FlashX within the broader GLM-4.7 release context, where GLM-4.7-Flash scored 59.2% on SWE-bench Verified and was benchmarked on consumer hardware via a 30B-parameter MoE architecture. FlashX is referenced as the higher-throughput sibling for users who need additional speed beyond the free tier.

Z.AI

Coverage

Z.AI (Zhipu AI) released GLM-4.7-FlashX as a high-speed, more cost-effective variant positioned alongside the GLM-4.7-Flash release, according to Z.AI's own social posts and developer documentation as paraphrased by the article. The article reports that GLM-4.7-Flash is targeted at local coding assistance and agentic workflows at the 30B level, with weights published on Hugging Face under the zai-org account and API access available via Z.ai. Use cases cited include creative writing, translation, long-context tasks, and role-playing beyond programming scenarios. Within the same launch, GLM-4.7-Flash is labeled as a free tier model on Z.ai with a default 1-concurrency limit, while GLM-4.7-FlashX is offered as an optional version aimed at higher-frequency call scenarios where faster response and lower per-call cost matter most. The article notes that real-world thresholds for "running locally" depend on deployment method and hardware, and that free-tier concurrency and commercial terms should be confirmed against the latest Z.ai pricing and terms page. Both variants are attributed to Z.AI as the creator.

Z.AI

Official sourceDocumentation

Z.AI's official developer documentation lists GLM-4.7-FlashX as one of three variants in the GLM-4.7 series, positioned as "Lightweight, High-Speed, and Affordable" alongside the full GLM-4.7 and GLM-4.7-Flash. The variant supports text input and output with a 200K context window and 128K maximum output tokens. The same overview page documents shared series capabilities available to FlashX, including thinking mode for scenario-based reasoning, real-time streaming output, function calling for external tool integration, context caching for long conversations, and structured output formats such as JSON. These capabilities position the model for agentic coding and complex task execution.

Videos about GLM-4.7-FlashX

More models around GLM-4.7-FlashX