Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Zhipu AI logo

Model details

GLM-5.3-FlashX

The model overview is temporarily unavailable.

Zhipu AIglm-5.3-flashxglm-flash

Quick Info

Powered by
Provider
Zhipu AI
Model key
glm-5.3-flashx
Release date
Sep 18, 2026
Last updated
Sep 18, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.37
Output token cost
$1.25

Limits

Output tokens
131,072 tokens
Context window
1,000,000 tokens

Transparent token rates

Compare GLM-5.3-FlashX pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GLM-5.3-FlashX

Zhipu AI

Coverage

BigGo Finance reports that GLM-5.3-FlashX was released on September 18, 2026, boosting inference speed from the prior 30-50 tokens per second to a peak of 200 tokens per second, a roughly five-to-six-fold improvement. The high-speed tier uses a distinct API identifier and is billed separately at approximately 2.5 times The speed gains are attributed to an inference cluster built on more than 100,000 Chinese-made AI accelerators, with InfraAgent, powered by GLM-5.3, participating in infrastructure optimization in what Zhipu describes as China's first publicly disclosed production recursive self-improvement deployment. The report also

Zhipu AI

Coverage

TechFlow Post confirms that Zhipu officially launched GLM-5.3-FlashX on September 18, 2026, with the API going live under the identifier "GLM-5.3-FlashX." The model had previously been accessible to global developers under the name "Ox Alpha." The report attributes the 200 tokens/s peak speed to inference computing power provided by 100,000 domestically produced chips, combined with further infrastructure investment and inference optimization by Zhipu. The notice is brief but corroborates the launch date and technical headline from other sources.

Zhipu AI

Coverage

APIMaster's technical write-up explicitly names GLM-5.3-FlashX, launched on September 18, 2026 with a published peak speed of 200 tokens/s and the API model ID "glm-5.3-flashx." The article clarifies that FlashX is not a new model but the same GLM-5.3-Flash system, sharing 320B total / 18B activated parameters, a 1M-to FlashX is priced at approximately $0.37 input and $1.25 output per 1M tokens, compared to $0.15 and $0.50 for the standard Flash tier, reflecting the throughput-optimized serving configuration. The piece notes that Flash is the open-weights variant included in the GLM Coding Plan with 3x quota, while FlashX is position

Zhipu AI

Coverage

Zhipu officially launched GLM-5.3-FlashX on September 18, 2026, as a high-speed serving tier built on the previously released GLM-5.3-Flash. According to the report, the new variant reaches a stated peak output speed of 200 tokens per second, targeting enterprise developers who need faster response times for coding and The article frames FlashX as Zhipu's first offering that combines intelligence, price, and speed simultaneously, and confirms that the GLM-5.3-FlashX API went live on the same day. The predecessor GLM-5.3-Flash had previously been distributed overseas under the name "Ox Alpha," and Zhipu credits recent infrastructure a

Zhipu AI

Coverage

The APIMaster blog index lists a dedicated comparison article titled "GLM-5.3-FlashX vs DeepSeek V4.1 Flash: Price, Speed, Limits, and Which to Use for Coding," published on September 18, 2026, alongside other Flash-tier coverage. The comparison positions both models as low-cost 1M-context open-weight Flash variants. The index also surfaces a dedicated GLM-5.3-FlashX guide covering its 200 tokens/s peak speed, model ID, pricing, context window, and benchmarks, published on the same launch date. This serves as a navigation entry to the deeper technical write-up rather than as standalone model evidence.

Videos about GLM-5.3-FlashX

More models around GLM-5.3-FlashX