Sulat.com
AI models
DevPass (LLM Gateway) logo

Model details

GLM-5.3-Flash

GLM-5.3-Flash is Z.ai's flagship efficient model in the GLM family, positioned as a 320B parameter sparse architecture that activates only 18B parameters per token. It is described as a multimodal open release that outperforms its predecessor GLM-5.2 across benchmarks and real-world workloads at a fraction of the cost, while approaching Claude Opus 4.8 on coding and agentic evaluations. The model is also known by the internal codename ox-alpha, reflecting Z.ai's continued iteration toward stronger reasoning and tool-using capabilities in a compact active footprint.

The model rests on a freshly trained base redesigned around a hybrid sparse and linear attention design, paired with Manifold-Constrained Hyper-Connections that improve scaling efficiency. This combination is intended to lower long-context serving costs without sacrificing accuracy, and it is pre-trained on a 30T-token multimodal corpus that spans text, visual, and other modalities. Practical fit centers on advanced reasoning and code generation, with a million-token context window and three selectable thinking modes (Low, High, and Max) so developers can dial effort up for complex agentic workflows or down for routine chat, making it attractive for teams that need frontier-style reasoning on a sparse, cost-efficient backbone.

DevPass (LLM Gateway)glm-5.3-flashglm

Quick Info

Powered by
Provider
DevPass (LLM Gateway)
Model key
glm-5.3-flash
Release date
Aug 26, 2026
Last updated
Aug 26, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.13
Output token cost
$0.40

Limits

Output tokens
131,072 tokens
Context window
1,048,576 tokens

Latest news about GLM-5.3-Flash

Videos about GLM-5.3-Flash

More models around GLM-5.3-Flash