Currently listed through these providers:
Model details
GLM-5.3-Flash
GLM-5.3-Flash is Z.ai's flagship efficient model in the GLM family, positioned as a 320B parameter sparse architecture that activates only 18B parameters per token. It is described as a multimodal open release that outperforms its predecessor GLM-5.2 across benchmarks and real-world workloads at a fraction of the cost, while approaching Claude Opus 4.8 on coding and agentic evaluations. The model is also known by the internal codename ox-alpha, reflecting Z.ai's continued iteration toward stronger reasoning and tool-using capabilities in a compact active footprint.
The model rests on a freshly trained base redesigned around a hybrid sparse and linear attention design, paired with Manifold-Constrained Hyper-Connections that improve scaling efficiency. This combination is intended to lower long-context serving costs without sacrificing accuracy, and it is pre-trained on a 30T-token multimodal corpus that spans text, visual, and other modalities. Practical fit centers on advanced reasoning and code generation, with a million-token context window and three selectable thinking modes (Low, High, and Max) so developers can dial effort up for complex agentic workflows or down for routine chat, making it attractive for teams that need frontier-style reasoning on a sparse, cost-efficient backbone.
Quick Info
Powered by- Provider
- DevPass (LLM Gateway)
- Model key
- glm-5.3-flash
- Release date
- Aug 26, 2026
- Last updated
- Aug 26, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.13
- Output token cost
- $0.40
Limits
- Output tokens
- 131,072 tokens
- Context window
- 1,048,576 tokens
Latest news about GLM-5.3-Flash
Videos about GLM-5.3-Flash
More models around GLM-5.3-Flash
This exact model name is also listed by 38 other providers.