Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Vercel AI Gateway logo

Model details

GLM 5.3 FlashX

GLM 5.3 FlashX is a high-speed inference variant released by Zhipu, operating under the developer brand Z.AI, and tailored for coding agents, real-time conversational interfaces, and long-running agentic workflows that benefit from rapid token generation. The model is positioned as a speed-focused sibling within the GLM-5.3 line, with its headline advantage being a maximum generation speed of up to 200 tokens per second. Compared with the standard GLM-5.3-Flash, FlashX is advertised as delivering up to 5× faster inference, with optimization concentrated at the inference infrastructure layer rather than in new architectural changes or benchmark gains. This makes it well suited for latency-sensitive applications such as interactive coding assistants and autonomous agent loops, where throughput per second directly shapes user experience.

FlashX accepts text and image inputs and produces text output, giving it a multimodal front end suitable for vision-grounded coding or document-aware agent tasks. The release was accompanied by infrastructure improvements attributed to the GLM-5.3-powered Infra Agent, which reportedly tripled end-to-end throughput on a deployment running across more than 100,000 domestic AI chips, signaling Zhipu's continued investment in scaling inference capacity rather than introducing new pre-training recipes. Practically, the model fits workloads that prioritize responsive, high-volume generation over deep one-shot reasoning, complementing rather than replacing denser GLM variants in a production stack.

Vercel AI Gatewayzai/glm-5.3-flashxglm-flash

Quick Info

Powered by
Provider
Vercel AI Gateway
Model key
zai/glm-5.3-flashx
Release date
Aug 26, 2026
Last updated
Aug 26, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.37
Output token cost
$1.25

Limits

Output tokens
131,072 tokens
Context window
1,000,000 tokens

Transparent token rates

Compare GLM 5.3 FlashX pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GLM 5.3 FlashX

Vercel AI Gateway

Coverage

Superpower Daily reports that Z.AI is giving Coding Plan subscribers three times the GLM-5.3 quota on the new GLM-5.3-Flash model, but is keeping the faster FlashX variant out of that subscription plan — both models are nevertheless available through Z.AI's API as `glm-5.3-flash` and `glm-5.3-flashx`. This access split The same article describes Flash as the first native multimodal model in the GLM-5 line, accepting video, images, text, and files with a one-million-token context window — a context length that also applies to FlashX. Z.AI reports the underlying model has 320B total parameters with 18B active per task, and claims 3.01×

Vercel AI Gateway

Coverage

Zhipu released GLM-5.3-FlashX on September 18, 2026, with an officially stated maximum inference speed of 200 tokens per second. The high-speed version is based on the previously released GLM-5.3-Flash, which had been introduced overseas under the name "Ox Alpha" and gained popularity among global developers for its in The API is live at launch, and the article explicitly names the GLM-5.3-FlashX variant and its 200 tokens/s peak speed. The coverage corroborates the launch facts from other independent outlets and confirms the model's intended audience of enterprise developers needing high-throughput, low-latency inference.

Vercel AI Gateway

Coverage

Zhipu released GLM-5.3-FlashX on September 18 as a high-speed variant of its GLM-5.3 series, raising inference output speed from 30–50 tokens per second to a maximum of 200 tokens per second, a roughly five-to-six-fold improvement. The FlashX tier uses a distinct model identifier ("GLM-5.3-FlashX") with separate billin The GLM-5.3-FlashX predecessor GLM-5.3-Flash previously launched anonymously as "Ox Alpha" on OpenCode and OpenRouter, logging over 62 trillion tokens in six days. The detailed pricing breakdown and technical differentiation from the base Flash model make this the most substantive independent coverage of the FlashX lau

Vercel AI Gateway

Coverage

Zhipu officially launched GLM-5.3-FlashX on September 18, 2026, with the API going live under the model identifier "GLM-5.3-FlashX." The model had previously been available to global developers under the name "Ox Alpha." According to the report, Zhipu leveraged inference computing power provided by 100,000 domestically The coverage confirms the FlashX variant name and its launch date explicitly, making it a direct and timely match for the GLM 5.3 FlashX subject. The 200 Tokens/s headline figure and the 100,000-chip inference cluster are both sourced to Zhipu's official statement. The piece is concise but serves as a primary announcem

Vercel AI Gateway

Coverage

Zhipu officially launched GLM-5.3-FlashX with the new version achieving an inference speed of up to 200 tokens/s, a fivefold increase over the existing GLM-5.3-Flash. Pricing has been raised to 2.5 times that of the original version. The launch reportedly drove Zhipu's afternoon gains to widen, with Z.AI shares rising While the piece is framed as a financial flash rather than a deep technical analysis, it explicitly confirms the GLM-5.3-FlashX launch name, the 200 tokens/s speed figure, the fivefold improvement over the standard Flash tier, and the 2.5x pricing multiplier. These facts are directly sourced to the launch announcement

Vercel AI Gateway

Coverage

AIBase reports that Zhipu AI officially launched GLM-5.3-FlashX on its BigModel platform on September 18, 2026, with the API going live at the same time. The model is identified as "GLM-5.3-FlashX," achieves a maximum output speed of 200 tokens/s, and is pitched around the three core strengths of intelligence, price, a The AIBase piece also claims — without independent confirmation — that Zhipu scaled to a domestic chip compute base of 100,000 units to meet rising demand for the Flash/X line, framing FlashX as a step forward in inference efficiency and domestic-LLM popularization. Developers can access the model through the official

Vercel AI Gateway

Coverage

GLM-5.3-FlashX is Zhipu's high-speed serving tier of GLM-5.3-Flash, released on September 18, 2026 with a published peak inference speed of 200 tokens/s. It shares the same underlying GLM-5.3-Flash system—320B total parameters with 18B activated, a 1M-token context window, up to 128,000 output tokens, and native image, FlashX is priced at approximately $0.37 input / $1.25 output per 1M tokens, compared to $0.15 / $0.50 for the standard Flash tier, and does not include open weights or the GLM Coding Plan's 3x quota benefit. The base GLM-5.3-Flash was open-sourced on August 26, 2026; FlashX is a separate inference configuration rather

Videos about GLM 5.3 FlashX

More models around GLM 5.3 FlashX