Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Z.AI logo

Model details

GLM-5.3-FlashX

GLM-5.3-FlashX is the streaming-focused tier of Z.AI's GLM Flash family, positioned as a multimodal coding model with a dedicated serving stack aimed at fast, interactive use. The model's emphasis is on responsive output for coding agents, tool-driven workflows, and applications where users watch tokens arrive in real time. Z.AI publishes a peak streaming rate of roughly 200 tokens per second for this tier, which the source material frames as the speed class that distinguishes GLM-5.3-FlashX from sibling GLM-5.3-Flash variants and that makes it relevant for repeated tool-loop calls where small per-turn waits add up.

As a Flash-tier open-weight model, GLM-5.3-FlashX is designed to balance cost and capability for production coding workloads. The combination of multimodal inputs and a large context window makes it suitable for tasks like reading entire repositories or documents and acting on them through tool calls. Practitioners who prioritize low-latency streaming, open weights for self-hosting, and multimodal reasoning inside agent loops will find GLM-5.3-FlashX a natural fit, while those who need lower cache-hit pricing or more flexible reasoning-effort controls may want to compare it against alternative Flash-tier offerings.

Z.AIglm-5.3-flashxglm-flash

Quick Info

Powered by
Provider
Z.AI
Model key
glm-5.3-flashx
Release date
Sep 18, 2026
Last updated
Sep 18, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.37
Output token cost
$1.25

Limits

Output tokens
131,072 tokens
Context window
1,000,000 tokens

Transparent token rates

Compare GLM-5.3-FlashX pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GLM-5.3-FlashX

Z.AI

Coverage

BigGo Finance's September 18, 2026 report independently corroborates Zhipu's launch of GLM-5.3-FlashX, describing it as the high-speed variant of the GLM-5.3 series that boosts inference speed from a prior 30–50 tokens/second to a maximum of 200 tokens/second — roughly a five-to-six-fold improvement. It states FlashX u The article adds infrastructure context that the speed gains are underpinned by an inference cluster built on more than 100,000 Chinese-made AI accelerators, with Zhipu's GLM team disclosing the cluster was built from scratch for the prior GLM generation. It also reports that InfraAgent, powered by GLM-5.3, participate

Z.AI

Coverage

Zhipu officially launched GLM-5.3-FlashX on September 18, 2026, with the model previously exposed to global developers under the name "Ox Alpha." Based on inference computing power provided by 100,000 domestically produced chips, Zhipu increased infrastructure investment and optimized inference performance, achieving s The official model identifier for the new tier is set to "GLM-5.3-FlashX," distinguishing it as a separately addressable API endpoint. Zhipu framed the release as combining intelligence, throughput, and cost-effectiveness in a single high-speed variant positioned for enterprise and developer use cases.

Z.AI

Coverage

Zhipu officially released GLM-5.3-FlashX on September 18, 2026, with an officially stated maximum inference speed of 200 tokens per second, and the API going live at launch. The high-speed variant is built on the previously released GLM-5.3-Flash (previously known externally as "Ox Alpha"), which gained popularity for Zhipu cited inference optimizations on infrastructure powered by 100,000 domestically produced chips as enabling the speed gains, and positioned FlashX as holding intelligence, price, and speed together simultaneously. The article frames the release as targeted at accelerating enterprise developer workflows requiring h

Z.AI

Coverage

The APIMaster recap dated September 18, 2026 explicitly states that GLM-5.3-FlashX is not a new model but a high-speed serving tier of GLM-5.3-Flash, sharing the same 320B-total/18B-activated MoE weights, 1M-token context, 128K output, and native image/video/file/text inputs. It documents the API model ID glm-5.3-flash The same page provides FlashX-specific launch facts — Zhipu announced it on September 18, 2026 with the API and official experience center both open at launch, and Zhipu reports speeds up to 200 tokens/s attributed to infrastructure work rather than a different checkpoint. It contrasts FlashX (designed for throughput,

Z.AI

Official sourceDocumentation

Z.AI's developer documentation confirms that GLM-5.3-FlashX is now live as a higher-throughput variant in the GLM-5.3 family, delivering inference speeds of 200 tokens per second for faster responses. It is part of the first native multimodal model line in the GLM-5 series, offering stronger intelligence than GLM-5.2 at low cost. The model supports a 1M-token context window with up to 128K maximum output tokens. Built on a hybrid sparse and linear attention architecture with 320B total and 18B activated parameters, GLM-5.3-FlashX reduces attention computation and KV cache by 3.01× and 4.44× versus GLM-5.3. It natively handles video, image, text, and file inputs for visual coding, Office workflows, and long-horizon agent tasks. The API model code is glm-5.3-flashx, accessible via Z.AI's Chat Completion API.

Z.AI

Coverage

The Hermes Agent integration guide confirms that FlashX is the same mixture-of-experts model as Flash — roughly 320 billion total parameters with about 18 billion active per token, natively multimodal (text, images, video in; text out), 1M-token context, and 128K maximum output — with Zhipu publishing no new benchmark On the developer-integration side, the page documents that on Z.AI's own API the model code is glm-5.3-flashx alongside glm-5.3-flash on the same documentation page, with the provider id "zai" and the environment variable GLM_API_KEY in Hermes. It notes that at publication time OpenRouter listed z-ai/glm-5.3-flash but

Videos about GLM-5.3-FlashX

More models around GLM-5.3-FlashX