Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
CoreWeave logo

Model details

GLM 5.3 Flash

GLM-5.3-Flash is Z.ai’s cost-focused, natively multimodal model for coding, automation, visual document work, and knowledge tasks. It uses a mixture-of-experts design with 320 billion total parameters and 18 billion active per token, trained on a 30-trillion-token multimodal corpus. Its hybrid linear-and-sparse attention architecture and Manifold-Constrained Hyper-Connections are intended to reduce the compute and memory demands of long-context processing.

The model is a strong fit for developers building coding agents and multi-step workflows, particularly when working across large codebases, charts, interfaces, or heterogeneous business documents. Reported results include 84.3 on Terminal-Bench 2.1, 63.4 on DeepSWE v1.1, and 78.4 on Toolathlon, while its visual evaluation results include 89.4 on CharXiv Reasoning with tools and 80.5 on MMVU. These results support its positioning as a capable, economical model for agentic coding and visually grounded professional tasks rather than a substitute for the strongest general-purpose models.

CoreWeavezai-org/GLM-5.3-Flashglm

Quick Info

Powered by
Provider
CoreWeave
Model key
zai-org/GLM-5.3-Flash
Release date
Aug 26, 2026
Last updated
Aug 26, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.15
Output token cost
$0.50

Limits

Output tokens
1,048,576 tokens
Context window
1,048,576 tokens

Latest news about GLM 5.3 Flash

Weights & Biases

CoverageBenchmark

GLM-5.3-Flash was released by Z.ai on August 26, 2026, under the MIT license with weights published on Hugging Face the same day. According to the page, it is a 320-billion-parameter mixture-of-experts model that activates only 18 billion parameters per token, making it the first natively multimodal model in the GLM-5 The article documents Z.ai's launch pricing of $0.15 per million input tokens and $0.50 per million output tokens, roughly a tenth of GLM 5.3's $1.40 and $4.40 pricing, and contrasts the model's benchmark posture against the still-unreleased GLM 5.3 (744B) weights. It highlights that while Flash posts competitive numbe

Weights & Biases

CoverageBenchmark

Before its identity was revealed, the model appeared on OpenRouter and OpenCode on August 20, 2026, listed only as "Ox Alpha" with no owner, no model card, a 1M-token context window, and text/image/video input. Within six days it became the most-used model on OpenRouter, processing roughly 23 trillion tokens — about 2. Weights landed on Hugging Face under an MIT license the same evening as the reveal, and Zhipu's Hong Kong shares closed more than 12% higher the next session, roughly 10x the January IPO price according to the article. Stripe had agreed to acquire OpenRouter the day before the model appeared, and Stripe CEO Patrick Col

Weights & Biases

CoverageBenchmark

Ampere's side-by-side comparison positions GLM-5.3-Flash as dramatically cheaper and natively multimodal relative to the GLM-5.3 flagship, while noting it delivers surprisingly close performance at far less compute. In independent Artificial Analysis testing, GLM-5.3 scores 60 on the Intelligence Index versus 57 for GL The article highlights an interesting asymmetry: GLM-5.3-Flash is not actually faster at generating tokens, with Artificial Analysis measuring GLM-5.3 at roughly 85 tokens per second versus about 49 tokens per second for Flash. For developers choosing between the two, it recommends GLM-5.3-Flash for most high-volume ap

Weights & Biases

CoverageBenchmark

DataCamp's coverage confirms GLM-5.3-Flash as Z.ai's cost-optimised sibling to the GLM-5.3 flagship, codenamed "Ox Alpha" during a stealth preview period before public reveal. The model uses a 320B-A18B mixture-of-experts configuration, supports a 1M-token context window, and is natively multimodal, distinguishing it f The article reports Z.ai's own benchmark of 84.3 on Terminal-Bench 2.1, placing GLM-5.3-Flash within striking distance of Claude Opus 4.8 (85.0) and GPT-5.6 Terra (87.4), and cites a blended API price of roughly $0.10 per million tokens compared to about $0.90 for GLM-5.3. It notes the practical deployment constraint t

Weights & Biases

Coverage

Z.ai has introduced GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series, released on August 26, 2026 under the MIT License with weights on Hugging Face. The model uses a mixture-of-experts design with 320B total parameters and 18B active per token, paired with a 1M-token context window and a hybrid s On benchmarks, GLM-5.3-Flash scores 57 on the Artificial Analysis Intelligence Index v4.1.1 at roughly $0.045 per task and approaches Claude Opus 4.8 on coding and agentic evaluations while outperforming GLM-5.2 across real-world workloads at about one-tenth the price. Standard Z.ai API pricing is $0.15 per 1M input to

Videos about GLM 5.3 Flash

More models around GLM 5.3 Flash