Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Nebius Token Factory logo

Model details

GLM-5.3-Flash

Z.ai built GLM-5.3-Flash as a newly trained, natively multimodal member of the GLM-5 series, rather than a lightweight continuation of an earlier checkpoint. Its sparse mixture-of-experts design contains 320 billion total parameters while activating 18 billion per token, combining linear and sparse attention to reduce the cost of long-context processing. Manifold-Constrained Hyper-Connections support scaling efficiency, and the model was pretrained on a 30-trillion-token multimodal corpus spanning text, images, and video.

The model is especially relevant to coding agents, visual coding loops, document and interface analysis, and tool-driven work. Its reported results show clear gains over GLM-5.2 across coding and agentic evaluations, including DeepSWE and AutomationBench, while approaching a leading closed model on a broader benchmark set. Native vision allows it to inspect rendered interfaces or other visual artifacts and refine outputs, although published results are not uniformly ahead of every comparator. The hybrid architecture, sparse activation, and publicly released weights make it a strong fit for teams seeking capable multimodal reasoning with more economical long-context operation.

Nebius Token Factoryzai-org/GLM-5.3-Flashglm-flash

Quick Info

Powered by
Provider
Nebius Token Factory
Model key
zai-org/GLM-5.3-Flash
Release date
Aug 26, 2026
Last updated
Aug 26, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.15
Output token cost
$0.50

Limits

Output tokens
1,024,000 tokens
Context window
1,024,000 tokens

Transparent token rates

Compare GLM-5.3-Flash pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GLM-5.3-Flash

Deep Infra

CoverageBenchmark

Z.ai launched GLM-5.3-Flash on August 26, 2026, as a 320-billion-parameter mixture-of-experts model that activates 18 billion parameters per token and ships with MIT-licensed weights on Hugging Face from day one. The supplied excerpt reproduces Z.ai's announcement text stating it is the first natively multimodal model The article frames GLM-5.3-Flash as not a distilled version of the flagship, noting it is built on a new base rather than post-trained from GLM-5.2. API pricing is reported at $0.15 per million input tokens and $0.50 per million output tokens, compared with $1.40 and $4.40 for GLM-5.3, and the author critiques Z.ai's b

Deep Infra

CoverageBenchmark

The cheapestinference technical profile (August 31, 2026) compiles GLM-5.3-Flash's headline specifications and applies an independent benchmark lens via Artificial Analysis: the model lands at 57 on the AA Intelligence Index at a blended $0.10 per million tokens ($0.09 per Index task) and sits on AA's intelligence-vs-c For context, the profile reports a 1M-token context window (up to 128K output), reasoning always on for the direct API, MIT-licensed weights on Hugging Face in BF16 and FP8, list pricing of $0.15 input / $0.50 output per million tokens (cached input $0.03) with a 50% launch promo through September 9, 2026, and measured

Deep Infra

Coverage

AI Intel Report's August 27, 2026 launch coverage corroborates the core GLM-5.3-Flash release facts with specific architecture numbers: Z.ai introduced the model on August 26, 2026 as the first natively multimodal entry in the GLM-5 series, equipped with 320 billion total parameters, 18 billion activated parameters, an The article also reconstructs the pre-launch "ox-alpha" testing phase, reporting that Z.ai deliberately deployed the model anonymously on OpenRouter and OpenCode to measure real-world usage without brand-influence bias, with traffic during that phase running entirely on Chinese AI chips to demonstrate ecosystem compati

Deep Infra

CoverageBenchmark

This comparison explicitly covers GLM-5.3-Flash alongside its sibling GLM-5.3. Per the supplied excerpt, GLM-5.3-Flash is a 320B-total/18B-active model from Z.ai with a 1M-token context, native multimodal input, and reasoning plus tool-use capabilities. It scores 57 on the Artificial Analysis Intelligence Index versus The piece argues that GLM-5.3-Flash is not a trimmed GLM-5.3 but a different model, and that despite a slower generation speed it is the better value for most high-volume applications and multimodal agents. It notes Flash is the only one of the two with publicly available weights, offering 3x usable quota on the Coding

Berget.AI

Coverage

Z.ai's alphaXiv paper introduces GLM-5.3-Flash as the first natively multimodal model in the GLM-5 series, a 320B-total / 18B-active Mixture-of-Experts model that combines sparse and linear attention in a hybrid architecture and adds Manifold-Constrained Hyper-Connections (mHC). It was trained on a 30T-token multimodal The paper documents standard API pricing of $0.15 per million input tokens, $0.50 per million output tokens, and $0.03 per million cached input tokens, and reports an Artificial Analysis Intelligence Index v4.1.1 score of 57 at $0.045 per task. Z.ai states the model outperforms GLM-5.2 across benchmarks and real-world

Deep Infra

CoverageDiscourse

A Hugging Face community PR opened on August 26, 2026 (and merged approximately six days later) by user SaylorTwift extracts evaluation results from the official zai-org/GLM-5.3-Flash model card's benchmark chart and adds them to the repository's .eval results/ directory. The PR records Terminal-Bench 2.1 at 84.3, Deep The PR is candid about provenance: the numbers were read visually from an embedded benchmark chart image (bench_53.png) in the README rather than reproduced via HF Jobs with inspect-ai, so no verification token accompanies them, and the cited arxiv:2602.15763 is identified as the original GLM-5 technical report from Fe

Deep Infra

Coverage

On August 26, 2026 at 7:42 PM, Z.ai (@Zai_org) publicly revealed GLM-5.3-Flash — the model previously appearing under the "ox-alpha" codename on OpenRouter/OpenCode — and the explainx.ai launch blog documents the reveal in near real time, citing the Z.ai tweet (740K+ views within an hour) and confirming that MIT-licens The explainx.ai piece also flags ecosystem context that supports the GLM-5.3-Flash release narrative: the stealth ox-alpha phase ran entirely on Chinese AI chips, the larger non-Flash GLM-5.3 missed its own August 28 open-weights staged-release target, and Unsloth shipped a Dynamic 3-bit GGUF sized to run on 128GB of R

Deep Infra

CoverageAnalysis

Z.ai released GLM-5.3-Flash on August 26, 2026, as the first natively multimodal model in the GLM-5 series. Per the supplied excerpt, it is a 320B-parameter mixture-of-experts model that activates only 18B parameters per token, runs natively in FP8, and supports a 1,048,576-token context window. The architecture uses h The deep dive situates GLM-5.3-Flash against GLM-4.6, GLM-5.2, and the GLM-5.3 flagship, noting it is built from a newly trained base rather than a post-train of GLM-5.2. It reports that GLM-5.3-Flash beats GLM-5.2 across six coding and agentic benchmarks at roughly one-tenth of the price and approaches Claude Opus 4.8

Nebius Token Factory

CoverageBenchmark

Together AI's listing describes GLM-5.3-Flash as Z.ai's first natively multimodal model in the GLM-5 series, built around frontier coding and agentic capability at flash-level inference cost. It is a 320B-parameter Mixture-of-Experts model with only 18B active parameters, introducing a hybrid architecture that combines The listing highlights four capability pillars: frontier coding at flash cost with gains well beyond GLM-5.2, visual intelligence that extends validation from functional correctness to the actual interfaces users see, efficient long context via IndexPool that cuts attention compute 3.0x and KV cache 4.4x versus GLM-5.3

Videos about GLM-5.3-Flash

More models around GLM-5.3-Flash