Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Privatemode AI logo

Model details

GLM-5.3-Flash

GLM-5.3-Flash is Z.ai's flagship open-weight release in the GLM-5 line, introduced as the series' first natively multimodal model. It carries 320B total parameters with only 18B active per token, a sparse-plus-linear attention hybrid that the team says is the first of its kind within the GLM series, designed to keep long-context serving inexpensive without sacrificing precision. Training was rebuilt around a 30T-token multimodal pre-training corpus, and Manifold-Constrained Hyper-Connections (mHC) were used to improve scaling efficiency, giving the model room to deliver stronger reasoning and coding behavior than its direct predecessor, GLM-5.2, while consuming a fraction of the compute and reaching roughly one-tenth the previous price point.

In real-world evaluations, GLM-5.3-Flash approaches Claude Opus 4.8 on coding and agentic benchmarks, a result the developers highlight as the model's defining strength. It also handles attachments, structured output, reasoning, tool calling, and image inputs alongside text generation. The release itself was unconventional: Z.ai tested the system anonymously as "ox-alpha" on OpenCode and OpenRouter before the public reveal, letting community traffic, including praise from Stripe's Patrick Collison, build the model's reputation ahead of launch. For practitioners, GLM-5.3-Flash fits projects that want frontier-class coding and agentic quality on locally deployable open weights, especially when long-context efficiency and multimodal grounding are priorities.

Privatemode AIglm-5.3-flashglm-flashbeta

Quick Info

Powered by
Provider
Privatemode AI
Model key
glm-5.3-flash
Release date
Aug 26, 2026
Last updated
Aug 26, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.8897
Output token cost
$4.4718

Limits

Output tokens
131,072 tokens
Context window
256,000 tokens

Transparent token rates

Compare GLM-5.3-Flash pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GLM-5.3-Flash

Privatemode AI

CoverageBenchmark

Z.ai released GLM-5.3-Flash on August 26, 2026, shipping MIT-licensed weights to Hugging Face on day one. According to the model announcement, it is a 320-billion-parameter mixture-of-experts model that activates only 18 billion parameters per token, built on a newly trained base rather than a post-train of GLM-5.2. Th The report identifies GLM-5.3-Flash as the first natively multimodal model in the GLM-5 series, with a one-million-token context window and a hybrid sparse-and-linear attention design. API pricing is set at $0.15 per million input tokens and $0.50 per million output tokens, roughly a tenth of the GLM-5.3 flagship rate.

Privatemode AI

Coverage

The guide confirms that Z.ai released GLM-5.3-Flash on August 26, 2026, as a 320B-total / 18B-active MoE with a one-million-token context window, native multimodality, and MIT-licensed weights. It documents the Ox Alpha backstory: an anonymous model appeared on a third-party API platform around August 20, was fingerpri The OpenRouter listing cited in the article specifies a context window of 1,310,720 tokens with up to 131,072 completion tokens, plus tool-calling and JSON output support — capabilities that matter for agent deployment. The piece also notes that the launch API price is roughly a tenth of the flagship GLM-5.3 rate, with

Privatemode AI

Coverage

The Medium-based technical overview describes GLM-5.3-Flash as Z.ai's August 26, 2026 release of a 320-billion-parameter MoE that activates approximately 18 billion parameters per token, with native multimodality, a one-million-token context window, and MIT-licensed open weights. The framing positions the model as deli The article reiterates the launch pricing of $0.15 per million input tokens and $0.50 per million output tokens and emphasizes the model's design priorities around compute-efficient intelligence. It frames the broader trend as moving beyond raw parameter counts toward maximizing usable capability per unit of inference

Privatemode AI

CoverageBenchmark

Independent Artificial Analysis testing cited in the comparison scores GLM-5.3-Flash at 57 on the Intelligence Index versus 60 for the full GLM-5.3 flagship. GLM-5.3-Flash is a 320B-total / 18B-active MoE with a one-million-token context window, native image input, and native multimodality — features the flagship lacks Pricing data shows GLM-5.3-Flash at $0.15 per million input tokens and $0.50 per million output tokens, roughly one-ninth the flagship's $1.40/$4.40 API rate, with a 3× usable Coding Plan quota advantage. The piece recommends GLM-5.3-Flash for most high-volume production workloads and multimodal agents, and GLM-5.3 for

Privatemode AI

CoverageAnalysis

The technical deep-dive confirms that Z.ai shipped GLM-5.3-Flash on August 26, 2026, describing it as a 320B-parameter MoE with 18B active per token, native FP8 inference, and a 1,048,576-token context window. It emphasizes that the model was trained from a newly designed base rather than derived from GLM-5.2's 744B ch The piece documents the Ox Alpha identity reveal, the MIT-licensed open-weight release, a public vLLM recipe, and an API price of $0.15/$0.50 per million tokens. It also reports claimed benchmark wins over GLM-5.2 across six coding and agentic tests and competitive long-horizon agent performance versus Claude Opus 4.8,

Videos about GLM-5.3-Flash

More models around GLM-5.3-Flash