Currently listed through these providers:
Model details
GLM-5.3-Flash
GLM-5.3-Flash is positioned as the first natively multimodal release in the GLM-5 lineup, designed to deliver stronger reasoning than GLM-5.2 while keeping serving costs low. Its hybrid architecture pairs 320B total parameters with only 18B activated parameters, mixing sparse attention and linear attention to cut attention compute and KV cache size by roughly 3× and 4.4× relative to GLM-5.3. Vision is integrated directly into the coding loop, letting the model observe interfaces, rendered output, and feedback as it iterates, which suits frontend, game, and 3D-style tasks where code and visual state need to stay in sync.
The model is aimed at practical coding and agentic work rather than just chat. Z.AI highlights its fit for professional workflows beyond coding, while the Ollama listing notes benchmark results approaching Claude Opus 4.8 on coding and agentic evaluations using just the 18B active parameters. A 1M-token context window supports long-running coding sessions and multi-step tool use, and availability through the GLM Coding Plan plus Ollama cloud runners such as Claude Code and OpenCode makes it straightforward to drop into existing developer setups.
Quick Info
Powered by- Provider
- Deep Infra
- Model key
- zai-org/GLM-5.3-Flash
- Release date
- Aug 26, 2026
- Last updated
- Aug 26, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.15
- Output token cost
- $0.50
Limits
- Output tokens
- 131,072 tokens
- Context window
- 1,048,576 tokens
Latest news about GLM-5.3-Flash
Videos about GLM-5.3-Flash
Recent tweets and retweets from Deep Infra
More models around GLM-5.3-Flash
This exact model name is also listed by 38 other providers.