Currently listed through these providers:
Model details
GLM-5.3 Flash (Consensus Protocol)
GLM-5.3-Flash is Z.ai's first natively multimodal model in the GLM-5 series, built to deliver frontier-class reasoning and coding ability at a fraction of the cost of larger predecessors. It is a sparse Mixture-of-Experts design with 320 billion total parameters but only 18 billion active per token, a deliberate step down from the previous generation's heavier backbone, and it is trained on a 30-trillion-token multimodal pre-training corpus that combines language, image, video, and document understanding into a single model. The creator positions it as outperforming its predecessor across benchmarks and real-world workloads while approaching top-tier closed models on coding and agentic evaluations, suggesting the model is aimed at developers and product teams that want strong tool use and reasoning without paying for a flagship-tier system.
What makes the architecture distinctive is a hybrid attention stack that mixes linear and standard attention to keep long-context serving cheap without sacrificing precision, combined with a multi-stream residual path and a native vision encoder that lets the same weights handle text and visual inputs. Independent analysis describes a roughly three-to-one pattern of efficient linear attention layers interleaved with multi-head latent and sparse attention layers, which is unusual among open frontier models and reflects a clear bet on efficiency over raw scale. Together these choices make GLM-5.3-Flash a practical fit for long-context assistants, coding agents, and multimodal applications where inference cost matters as much as peak capability.
Quick Info
Powered by- Provider
- LLM Gateway
- Model key
- consensusprotocol/glm-5.3-flash
- Release date
- Aug 26, 2026
- Last updated
- Aug 26, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.10
- Output token cost
- $0.25
Limits
- Output tokens
- 131,072 tokens
- Context window
- 1,048,576 tokens
Transparent token rates
Compare GLM-5.3 Flash (Consensus Protocol) pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about GLM-5.3 Flash (Consensus Protocol)
No articles yet. Fetch the latest news to show it here.
Videos about GLM-5.3 Flash (Consensus Protocol)
More models around GLM-5.3 Flash (Consensus Protocol)
This exact model name is also listed by 60 other providers.