Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Melious logo

Model details

GLM-5.3-Flash

GLM-5.3-Flash is a Mixture-of-Experts model from Z.ai and the first natively multimodal release in the GLM-5 series, accepting text, image, and video inputs within a one-the cataloged API limit. It carries 320B total parameters with 18B active per token, a sparse design that aims to deliver frontier-class capability while keeping the compute required per inference comparatively modest. The model first reached the public under the anonymous "stealth/ox-alpha" listing, and Z.ai confirmed on August 26, 2026 that this stealth entry was, in fact, GLM-5.3-Flash, marking it as a new chapter in the GLM family rather than a refresh of an older checkpoint.

Beyond raw capability, GLM-5.3-Flash is positioned for practical, local-friendly deployment. Independent reporting describes it as the first open multimodal model that runs on a single Mac and was served without NVIDIA hardware, suggesting the 18B-active MoE layout was tuned for efficient local and alternative-accelerator inference rather than only hyperscale data centers. That combination of native multimodality, an exceptionally long context, and a sparse activation pattern makes the model a strong fit for builders who want to mix long-document reasoning with image or video understanding in a single call, particularly where open weights and self-hosting matter.

Meliousglm-5.3-flashglm-flash

Quick Info

Powered by
Provider
Melious
Model key
glm-5.3-flash
Release date
Aug 26, 2026
Last updated
Aug 26, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.11592
Output token cost
$0.46368

Limits

Output tokens
131,072 tokens
Context window
1,000,000 tokens

Transparent token rates

Compare GLM-5.3-Flash pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GLM-5.3-Flash

Melious

CoverageBenchmark

Z.ai (also known as Zhipu AI) released GLM-5.3-Flash on August 26, 2026, with 320 billion total parameters and 18 billion active per token, making it the first natively multimodal model in the GLM-5 series with a 1M-token context window. Weights landed on Hugging Face under an MIT license that same evening, and Zhipu's For six days before launch, an anonymous listing called "Ox Alpha" topped OpenRouter's usage charts, processing about 23 trillion tokens, the biggest launch in OpenRouter's history at roughly 2.3 times the volume of the next model, and ending DeepSeek's 56-day run at the top on OpenCode. On August 26, Bloomberg confirm

Melious

CoverageBenchmark

FelloAI's piece confirms GLM-5.3-Flash shipped on August 26, 2026 with MIT-licensed weights on Hugging Face on day one, and details it as a 320B-total, 18B-active mixture-of-experts model built on a newly trained base rather than a post-train of GLM 5.2. The article notes the model is natively multimodal with a one-mil Beyond specs and pricing, FelloAI identifies GLM-5.3-Flash as the previously anonymous Ox Alpha model that had been serving traffic on OpenRouter, and points out that Z.ai's Hugging Face organisation has no GLM 5.3 repository at all, so the 744B flagship weights remain missing as of the September 1, 2026 update. The pi

Melious

Coverage

Z.ai released GLM-5.3-Flash on August 26, 2026, positioning the model as a 320-billion-parameter mixture-of-experts design that activates roughly 18 billion parameters per token, according to the supplied Medium write-up. The piece reports that the model supports multimodal inputs, offers a context window of up to one The Medium article is a third-party analysis rather than a Z.ai announcement, and it overlaps significantly with the more technical FelloAI and Ampere.sh comparisons on parameter counts, context length, and license. It remains useful as a dated, model-focused summary of the GLM-5.3-Flash launch, with the model explicit

Melious

CoverageBenchmark

Ampere.sh's comparison explicitly names GLM 5.3 Flash as a 320B-total, 18B-active Z.ai model with a one-million-token context, native multimodal input, MIT open weights, and reasoning and tool-use support. It cites Artificial Analysis intelligence scores of 60 for GLM 5.3 versus 57 for GLM 5.3 Flash, output speeds of r The post frames GLM 5.3 as the stronger choice for hard coding, reasoning, and long-horizon engineering, while GLM 5.3 Flash is positioned as dramatically cheaper with closer performance and the only option with native image input and multimodal capability. Although it is a third-party comparison rather than an officia

Melious

CoverageDiscourse

The Hugging Face discussion is a community PR that adds evaluation results to the zai-org/GLM-5.3-Flash repository, reporting Terminal-Bench 2.1 at 84.3, DeepSWE at 63.4, and Humanity's Last Exam at 55.3, with skipped entries for Agent's Last Exam (26.3), AutomationBench v1.0.6 (48.8), and GDPval-AA v2 (1773) because n The thread also flags that the model card cites arxiv:2602.15763 as the GLM-5 technical report, but the original February 2026 paper predates GLM-5.3-Flash and does not contain its specific numbers, so it was not used as the source. This is repository-process news rather than a model announcement, but it provides versi

Melious

CoverageAnalysis

Z.ai released GLM-5.3-Flash on August 26, 2026, as a 320-billion-parameter mixture-of-experts model that activates only 18 billion parameters per token, runs natively in FP8, and supports a 1,048,576-token context window. It is the first natively multimodal model in the GLM-5 series, combining text, image, and video un Before launch, the model ran anonymously on OpenRouter for about six days as "Ox Alpha," processing roughly 23 trillion tokens before Bloomberg and Business confirmed its identity as GLM-5.3-Flash from Z.ai. The API is priced at $0.15 per million input tokens and $0.50 per million output tokens, and on the Artificial A

Videos about GLM-5.3-Flash

More models around GLM-5.3-Flash