Sulat.com
AI models
LLM Gateway logo

Model details

GLM-5.3 Flash (Z AI)

GLM-5.3 Flash is presented as a new foundation within the GLM-5 line, built on a different base from its predecessor and described in coverage as a mixture-of-experts system with 320 billion total parameters and about 18 billion active per token. This design concentrates computation on a smaller subset of parameters for each token, suggesting a practical balance between broad model capacity and more selective inference work.

The model is positioned for demanding text applications such as reasoning, coding, and long-context tasks, with the anonymous Ox Alpha service having drawn attention for its extensive context window before a Nebius disclosure linked that endpoint to GLM-5.3 Flash. Its sparse activation pattern may appeal to teams that want high-capacity behavior without activating the full parameter set at every step, while the supplied evidence does not establish detailed training lineage or independently verified benchmark gains.

LLM Gatewayzai/glm-5.3-flashglm

Quick Info

Powered by
Provider
LLM Gateway
Model key
zai/glm-5.3-flash
Release date
Aug 26, 2026
Last updated
Aug 26, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.15
Output token cost
$0.50

Limits

Output tokens
131,072 tokens
Context window
1,048,576 tokens

Latest news about GLM-5.3 Flash (Z AI)

Videos about GLM-5.3 Flash (Z AI)

More models around GLM-5.3 Flash (Z AI)