Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Ofox logo

Model details

GLM-5.3-FlashX

The model overview is being prepared.

Ofoxz-ai/glm-5.3-flashxglm-flash

Quick Info

Powered by
Provider
Ofox
Model key
z-ai/glm-5.3-flashx
Release date
Sep 18, 2026
Last updated
Sep 18, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.15
Output token cost
$0.50

Limits

Output tokens
131,072 tokens
Context window
1,000,000 tokens

Transparent token rates

Compare GLM-5.3-FlashX pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GLM-5.3-FlashX

Ofox

Coverage

TechFlow's September 18, 2026 newsletter reports Zhipu's official launch of GLM-5.3-FlashX, confirming the model was previously previewed to global developers under the codename "Ox Alpha." Zhipu stated the 200 tokens/s speed was achieved through increased infrastructure investment and inference optimization on a computing cluster backed by 100,000 domestically produced chips, with the FlashX API going live under the identifier "GLM-5.3-FlashX." The launch framing positions FlashX as combining intelligence, price, and speed simultaneously for the first time within the GLM-5 series. The article notes rapid prior adoption of the underlying GLM-5.3-Flash among global developers drove Zhipu to scale inference capacity, making FlashX a direct response to enterprise demand for higher-throughput, lower-latency access to the same model weights.

Ofox

Official sourceDocumentation

Z.ai's official developer documentation introduces GLM-5.3-FlashX as the high-speed variant of GLM-5.3-Flash, achieving inference speeds of up to 200 tokens/s on the same multimodal architecture. The model shares Flash's 320B total / 18B active parameters, hybrid sparse-plus-linear attention, 1M-token context window, and 128K maximum output, with API code "glm-5.3-flashx" available via the Chat Completion API. No new benchmarks were published because weights remain unchanged from Flash. The documentation positions FlashX alongside Flash on the GLM Coding Plan, emphasizing native multimodal visual coding, professional workflow support (Office, financial research, document processing), and recommended sampling settings of temperature 1, top_p 0.95, and reasoning_effort max. The FlashX-specific speed gain is delivered through inference optimization on Zhipu's domestic-chip infrastructure rather than capability changes, making it a serving-layer upgrade targeted at high-throughput, low-latency enterprise usage.

Videos about GLM-5.3-FlashX

More models around GLM-5.3-FlashX