Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
OpenRouter logo

Model details

GLM 5.3 FlashX

GLM 5.3 FlashX is Z.AI's high-speed inference variant within the GLM-5.3 family, designed for coding agents, real-time interactions, and long-running agentic workloads. It belongs to the first native multimodal generation of the GLM-5 series, delivering what Z.AI describes as stronger intelligence than the preceding GLM-5.2 line while keeping cost and latency low. Compared with the base GLM-5.3-Flash, FlashX is reported to run up to five times faster, reaching roughly 200 tokens per second, with a clear focus on low-latency and high-throughput serving.

Underneath, the family uses a hybrid architecture with around 320 billion total parameters and about 18 billion activated per token, combining sparse and linear attention to trim compute and memory pressure. Z.AI's documentation frames this as a frontier-level design that cuts attention computation by roughly threefold and KV cache by more than fourfold relative to GLM-5.3, helping preserve long-context quality. Native multimodal visual coding sits inside the agent loop, letting the model observe interfaces, rendered results, and interaction feedback to drive coding, browser, and GUI tasks, while also extending into office work, document handling, and financial research with structured deliverables.

OpenRouterz-ai/glm-5.3-flashxglm

Quick Info

Powered by
Provider
OpenRouter
Model key
z-ai/glm-5.3-flashx
Release date
Sep 18, 2026
Last updated
Sep 18, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.37
Output token cost
$1.25

Limits

Output tokens
131,072 tokens
Context window
1,048,576 tokens

Transparent token rates

Compare GLM 5.3 FlashX pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GLM 5.3 FlashX

OpenRouter

Coverage

Superpower Daily reported on September 18, 2026 that Z.AI added GLM-5.3-Flash to its Coding Plan with three times the GLM-5.3 quota, but deliberately excluded the faster FlashX variant from the subscription for now. Both models are available through Z.AI's API under the identifiers glm-5.3-flash and glm-5.3-flashx resp The same article describes GLM-5.3-Flash as the first native multimodal model in the GLM-5 line, accepting video, images, text, and files with a one-million-token context window, 320B total / 18B active parameters, and substantially lower attention and KV-cache costs on long prompts compared with GLM-5.3. FlashX's adve

OpenRouter

Coverage

Zhipu released GLM-5.3-FlashX on September 18, 2026 as a high-speed serving tier of the GLM-5.3-Flash model, with the vendor officially stating a peak inference speed of up to 200 tokens per second. The tier builds on the previously released GLM-5.3-Flash (which had circulated globally among developers under the name " According to the same coverage, FlashX is positioned to deliver intelligence, price, and speed simultaneously for enterprise developers, with the API going live alongside the announcement. The article frames the speed boost as the result of continued infrastructure investment and inference optimization on Chinese-made

OpenRouter

Coverage

BigGo Finance detailed Zhipu's September 18, 2026 launch of GLM-5.3-FlashX, reporting that inference speed rose from roughly 30–50 tokens/s on the standard Flash tier to a stated maximum of 200 tokens/s — a five-to-six-fold improvement that the outlet says closes in on DeepSeek's published performance. The high-speed t The same report describes an inference cluster built on more than 100,000 Chinese-made AI accelerators and credits Zhipu's InfraAgent (powered by GLM-5.3) with participating in infrastructure optimization, framed by Zhipu as China's first publicly disclosed recursive self-improvement deployment in production. It also n

OpenRouter

Coverage

TechFlow reported on September 18, 2026 that Zhipu officially launched GLM-5.3-FlashX, with the API going live under the explicit model identifier "GLM-5.3-FlashX". The piece notes the model had previously been available to global developers under the "Ox Alpha" name and confirms Zhipu's stated peak speed of up to 200 The report attributes both the launch and the infrastructure description directly to Zhipu, distinguishing the creator from OpenRouter or any routing layer. The brief corroborates the model identifier string, the launch date, and the speed claim that other same-day sources also report, providing timely confirmation of

OpenRouter

Coverage

APIMaster's technical breakdown distinguishes GLM-5.3-FlashX from its base sibling: FlashX is the speed-optimized serving tier of the same GLM-5.3-Flash checkpoint (320B total / 18B activated parameters, 1M-token context, 128K max output, native video/image/text/file multimodal inputs, text output), with the API model The article cites Zhipu's list pricing of $0.37 input and $1.25 output per 1M tokens for FlashX against $0.15 and $0.50 for the base Flash tier, and notes that Flash is the open-weights variant included in the GLM Coding Plan (with 3× the GLM-5.3 quota), while FlashX is positioned for throughput-sensitive, low-latency,

Videos about GLM 5.3 FlashX

More models around GLM 5.3 FlashX