Currently listed through these providers:
Model details
GLM-5.3 Flash (Z AI)
GLM-5.3 Flash is presented as a new foundation within the GLM-5 line, built on a different base from its predecessor and described in coverage as a mixture-of-experts system with 320 billion total parameters and about 18 billion active per token. This design concentrates computation on a smaller subset of parameters for each token, suggesting a practical balance between broad model capacity and more selective inference work.
The model is positioned for demanding text applications such as reasoning, coding, and long-context tasks, with the anonymous Ox Alpha service having drawn attention for its extensive context window before a Nebius disclosure linked that endpoint to GLM-5.3 Flash. Its sparse activation pattern may appeal to teams that want high-capacity behavior without activating the full parameter set at every step, while the supplied evidence does not establish detailed training lineage or independently verified benchmark gains.
Quick Info
Powered by- Provider
- LLM Gateway
- Model key
- zai/glm-5.3-flash
- Release date
- Aug 26, 2026
- Last updated
- Aug 26, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.15
- Output token cost
- $0.50
Limits
- Output tokens
- 131,072 tokens
- Context window
- 1,048,576 tokens
Latest news about GLM-5.3 Flash (Z AI)
Videos about GLM-5.3 Flash (Z AI)
More models around GLM-5.3 Flash (Z AI)
This exact model name is also listed by 38 other providers.