Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
SiliconFlow logo

Model details

GLM-5.3-Flash

Z.ai built GLM-5.3-Flash as a newly trained, natively multimodal member of the GLM-5 series, rather than a lightweight continuation of an earlier checkpoint. Its sparse mixture-of-experts design contains 320 billion total parameters while activating 18 billion per token, combining linear and sparse attention to reduce the cost of long-context processing. Manifold-Constrained Hyper-Connections support scaling efficiency, and the model was pretrained on a 30-trillion-token multimodal corpus spanning text, images, and video.

The model is especially relevant to coding agents, visual coding loops, document and interface analysis, and tool-driven work. Its reported results show clear gains over GLM-5.2 across coding and agentic evaluations, including DeepSWE and AutomationBench, while approaching a leading closed model on a broader benchmark set. Native vision allows it to inspect rendered interfaces or other visual artifacts and refine outputs, although published results are not uniformly ahead of every comparator. The hybrid architecture, sparse activation, and publicly released weights make it a strong fit for teams seeking capable multimodal reasoning with more economical long-context operation.

SiliconFlowzai-org/GLM-5.3-Flashglm-flash

Quick Info

Powered by
Provider
SiliconFlow
Model key
zai-org/GLM-5.3-Flash
Release date
Aug 26, 2026
Last updated
Aug 26, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.15
Output token cost
$0.50

Limits

Output tokens
262,000 tokens
Context window
1,049,000 tokens

Transparent token rates

Compare GLM-5.3-Flash pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GLM-5.3-Flash

RunInfra

CoverageBenchmark

An anonymous listing called "Ox Alpha" appeared on OpenRouter and OpenCode on August 20, 2026, offering free access with a 1M-token context window and text, image, and video input but no listed owner or model card. Within six days it became the most-used model on OpenRouter, processing about 23 trillion tokens, the lar GLM-5.3-Flash ships as a 320B-total / 18B-active MoE, the first natively multimodal model in the GLM-5 series, and the cheapest model at its score on Artificial Analysis's Intelligence Index. Weights landed on Hugging Face under an MIT license the same evening the identity was confirmed, and Zhipu's Hong Kong shares cl

Deep Infra

CoverageBenchmark

Z.ai launched GLM-5.3-Flash on August 26, 2026, as a 320-billion-parameter mixture-of-experts model that activates 18 billion parameters per token and ships with MIT-licensed weights on Hugging Face from day one. The supplied excerpt reproduces Z.ai's announcement text stating it is the first natively multimodal model The article frames GLM-5.3-Flash as not a distilled version of the flagship, noting it is built on a new base rather than post-trained from GLM-5.2. API pricing is reported at $0.15 per million input tokens and $0.50 per million output tokens, compared with $1.40 and $4.40 for GLM-5.3, and the author critiques Z.ai's b

RunInfra

Coverage

GLM-5.3-Flash became available for free access in two ways following its August 26, 2026 release by Z.ai: downloading the MIT-licensed open weights for self-hosting, or riding launch discounts that cut the already-low API price in half. The model is a mixture-of-experts design with 320 billion total parameters and 18 b Per the OpenRouter listing and release notes, GLM-5.3-Flash has a 1,310,720-token context window with up to 131,072 completion tokens, and supports tool calling and JSON-formatted outputs, the two features that matter most for running it inside an agent rather than a chat window. Z.ai positions the model for efficient

Deep Infra

CoverageBenchmark

The cheapestinference technical profile (August 31, 2026) compiles GLM-5.3-Flash's headline specifications and applies an independent benchmark lens via Artificial Analysis: the model lands at 57 on the AA Intelligence Index at a blended $0.10 per million tokens ($0.09 per Index task) and sits on AA's intelligence-vs-c For context, the profile reports a 1M-token context window (up to 128K output), reasoning always on for the direct API, MIT-licensed weights on Hugging Face in BF16 and FP8, list pricing of $0.15 input / $0.50 output per million tokens (cached input $0.03) with a 50% launch promo through September 9, 2026, and measured

RunInfra

CoverageBenchmark

An independent head-to-head comparison using Artificial Analysis data puts GLM-5.3-Flash at an Intelligence Index score of 57 versus 60 for the flagship GLM-5.3, but Flash costs roughly one-ninth as much at normal API pricing ($0.15/$0.50 per million tokens versus $1.40/$4.40) and activates only 18B parameters versus 4 Per the comparison, GLM-5.3 wins on overall intelligence, complex coding, long-horizon engineering, terminal coding, and raw output speed, while GLM-5.3-Flash wins on API price, cost-performance, image understanding, visual coding, browser agents, computer-use agents, and Coding Plan quota (3x usable quota versus 1x re

Deep Infra

Coverage

AI Intel Report's August 27, 2026 launch coverage corroborates the core GLM-5.3-Flash release facts with specific architecture numbers: Z.ai introduced the model on August 26, 2026 as the first natively multimodal entry in the GLM-5 series, equipped with 320 billion total parameters, 18 billion activated parameters, an The article also reconstructs the pre-launch "ox-alpha" testing phase, reporting that Z.ai deliberately deployed the model anonymously on OpenRouter and OpenCode to measure real-world usage without brand-influence bias, with traffic during that phase running entirely on Chinese AI chips to demonstrate ecosystem compati

Berget.AI

Coverage

Z.ai's alphaXiv paper introduces GLM-5.3-Flash as the first natively multimodal model in the GLM-5 series, a 320B-total / 18B-active Mixture-of-Experts model that combines sparse and linear attention in a hybrid architecture and adds Manifold-Constrained Hyper-Connections (mHC). It was trained on a 30T-token multimodal The paper documents standard API pricing of $0.15 per million input tokens, $0.50 per million output tokens, and $0.03 per million cached input tokens, and reports an Artificial Analysis Intelligence Index v4.1.1 score of 57 at $0.045 per task. Z.ai states the model outperforms GLM-5.2 across benchmarks and real-world

Deep Infra

CoverageDiscourse

A Hugging Face community PR opened on August 26, 2026 (and merged approximately six days later) by user SaylorTwift extracts evaluation results from the official zai-org/GLM-5.3-Flash model card's benchmark chart and adds them to the repository's .eval results/ directory. The PR records Terminal-Bench 2.1 at 84.3, Deep The PR is candid about provenance: the numbers were read visually from an embedded benchmark chart image (bench_53.png) in the README rather than reproduced via HF Jobs with inspect-ai, so no verification token accompanies them, and the cited arxiv:2602.15763 is identified as the original GLM-5 technical report from Fe

Deep Infra

Coverage

On August 26, 2026 at 7:42 PM, Z.ai (@Zai_org) publicly revealed GLM-5.3-Flash — the model previously appearing under the "ox-alpha" codename on OpenRouter/OpenCode — and the explainx.ai launch blog documents the reveal in near real time, citing the Z.ai tweet (740K+ views within an hour) and confirming that MIT-licens The explainx.ai piece also flags ecosystem context that supports the GLM-5.3-Flash release narrative: the stealth ox-alpha phase ran entirely on Chinese AI chips, the larger non-Flash GLM-5.3 missed its own August 28 open-weights staged-release target, and Unsloth shipped a Dynamic 3-bit GGUF sized to run on 128GB of R

Deep Infra

CoverageAnalysis

Z.ai released GLM-5.3-Flash on August 26, 2026, as the first natively multimodal model in the GLM-5 series. Per the supplied excerpt, it is a 320B-parameter mixture-of-experts model that activates only 18B parameters per token, runs natively in FP8, and supports a 1,048,576-token context window. The architecture uses h The deep dive situates GLM-5.3-Flash against GLM-4.6, GLM-5.2, and the GLM-5.3 flagship, noting it is built from a newly trained base rather than a post-train of GLM-5.2. It reports that GLM-5.3-Flash beats GLM-5.2 across six coding and agentic benchmarks at roughly one-tenth of the price and approaches Claude Opus 4.8

Videos about GLM-5.3-Flash

More models around GLM-5.3-Flash