Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
OpenCode Zen logo

Model details

GLM-5.3-Flash

GLM-5.3-Flash is positioned as the first natively multimodal entry in the GLM-5 lineup, built by Z.ai around a redesigned base model that blends sparse and linear attention in a single hybrid architecture. That design, paired with Manifold-Constrained Hyper-Connections, is intended to cut long-context serving cost without sacrificing retrieval-style precision, and it is trained on a fresh 30T-token multimodal corpus that replaces the recipe used for earlier GLM-5 variants. The model carries 320B total parameters but only activates about 18B at inference time, giving it a notably favorable ratio of capability to compute compared with its predecessors in the family.

In practice, GLM-5.3-Flash is tuned for coding and agentic workloads where long horizons and tool use matter. Z.ai reports that it surpasses GLM-5.2 across internal benchmarks and real-world evaluations while approaching the coding and agentic performance of Claude Opus 4.8, all at a fraction of the serving cost. The model was stress-tested anonymously as "Ox Alpha" on OpenCode and OpenRouter, where it became the most popular model of the week, with traffic served entirely on Chinese AI chips. It is also featured as a curated option on the $10-per-month OpenCode Go subscription tier, making it an accessible choice for developers who want frontier-flavored reasoning without paying flagship-tier prices.

OpenCode Zenglm-5.3-flashglm-flash

Quick Info

Powered by
Provider
OpenCode Zen
Model key
glm-5.3-flash
Release date
Aug 26, 2026
Last updated
Aug 26, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.15
Output token cost
$0.50

Limits

Output tokens
131,072 tokens
Context window
1,000,000 tokens

Transparent token rates

Compare GLM-5.3-Flash pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GLM-5.3-Flash

OpenCode Zen

Official sourceDocumentation

OpenCode's own documentation for the "OpenCode Go" subscription tier explicitly lists GLM-5.3-Flash as one of the curated open coding models accessible through the $10/month plan. The page describes Go as a low-cost, internationally focused subscription that sits alongside OpenCode Zen, providing stable global access t For developers, setup works through the standard OpenCode provider flow: sign in to OpenCode Zen, subscribe to Go, copy the generated API key, then run the /connect command in the OpenCode TUI, select OpenCode Go, and paste the key. Running /models surfaces the available Go lineup, and traffic is monitored to prevent a

OpenCode Zen

Coverage

The Medium article by Greek Ai summarizes Z.ai's release of GLM-5.3-Flash on August 26, 2026, describing it as the newest member of the GLM-5 family and the first natively multimodal model in the series. The supplied excerpt states it has 320 billion total parameters with roughly 18 billion active parameters, supports The excerpt also documents the pre-launch period in which a model called "Ox Alpha" appeared on platforms including OpenRouter and OpenCode before Z.ai officially confirmed it as an anonymous preview of GLM-5.3-Flash, a detail the article uses to motivate the model's real-world adoption prior to its public identity bei

OpenCode Zen

CoverageBenchmark

Ampere's comparison blog contrasts GLM 5.3 with the subject model GLM 5.3 Flash on August 27, 2026, citing Artificial Analysis figures. GLM 5.3 Flash scores 57 on the Intelligence Index versus 60 for the flagship GLM 5.3, but costs roughly one-ninth as much, activates only 18B parameters against 40B, and is the first n The page positions GLM 5.3 Flash as the better value for high-volume applications and multimodal or browser-based agents, where it wins categories including image understanding, visual coding, browser agents, and computer-use agents, and offers three times the usable quota under the Z.ai Coding Plan. GLM 5.3 remains th

OpenCode Zen

CoverageAnalysis

Local AI Zone's technical deep dive documents GLM 5.3 Flash as shipped by Z.ai on August 26, 2026: a 320B-total, 18B-active mixture-of-experts model running natively in FP8, with a 1,048,576-token context window, hybrid sparse and linear attention, and native multimodality (text, image, and video input, text output). I The page reveals that GLM 5.3 Flash is the model previously previewed anonymously as "Ox Alpha" on OpenRouter before launch, and reports that it beats GLM 5.2 across six coding and agentic benchmarks while approaching Claude Opus 4.8 on long-horizon agent tasks at roughly one-tenth the price of Z.ai's flagship. The fla

OpenCode Zen

CoverageBenchmark

Artificial Analysis' evaluation page for GLM-5.3-Flash records the model as a Z.ai open-weight release with 320 billion total parameters, 18 billion active per token, a 1-million-token context window, and text-and-image input producing text output. The supplied excerpt shows the reasoning-enabled variant scoring 46 on On quantitative performance, the page reports output throughput of 48.5 tokens per second (51st of 112 models, 48 tok/s rounded in the summary, scoring 1 of 4 speed units) and pricing of $0.15 per million input tokens and $0.50 per million output tokens with an 83% cache discount, yielding a cost-per-Intelligence-Index

Videos about GLM-5.3-Flash

More models around GLM-5.3-Flash