Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
302.AI logo

Model details

GLM-5.3-Flash

GLM-5.3-Flash is a cost-optimized member of Z.ai's GLM model family, positioned as a lighter sibling to the flagship GLM-5.3. It first drew community attention under an anonymous preview codename before being publicly revealed as "Ox Alpha," and community discussion threads about its parameter layout appeared on the NVIDIA Developer Forums on the same day as the broader launch. The model's design centers on a mixture-of-experts architecture with 320B aggregate parameters and 18B active per token, a configuration that aims to keep inference economical while preserving broad capability for downstream tasks.

The model is natively multimodal and supports the cataloged API limit token context window, making it suited to long-form reasoning, multi-document workflows, and agentic pipelines where extended context retention matters. On Terminal-Bench 2.1, GLM-5.3-Flash scores 84.3, placing it close to Claude Opus 4.8 at 85.0 and GPT-5.6 Terra at 87.4, which the reporting source frames as near-frontier coding and agentic performance at a lower cost tier. Practically, this combination of MoE efficiency, large context, and competitive coding scores makes GLM-5.3-Flash a fit for teams building budget-conscious assistants, code-generation tools, and multi-step automated workflows that need strong tool use without flagship-tier expense.

302.AIglm-5.3-flashglm-flash

Quick Info

Powered by
Provider
302.AI
Model key
glm-5.3-flash
Release date
Aug 26, 2026
Last updated
Aug 26, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.075
Output token cost
$0.25

Limits

Output tokens
131,072 tokens
Context window
1,000,000 tokens

Transparent token rates

Compare GLM-5.3-Flash pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GLM-5.3-Flash

302.AI

CoverageBenchmark

Z.ai released GLM 5.3 Flash on August 26, 2026, publishing MIT-licensed weights on Hugging Face on day one, and the article confirms the model is identical to "Ox Alpha," the anonymous listing that had been quietly serving traffic on OpenRouter. The model is a 320-billion-parameter mixture-of-experts with 18 billion pa The piece highlights Z.ai's published pricing of $0.15 per million input tokens and $0.50 per million output tokens, compared with $1.40 and $4.40 respectively for GLM 5.3, and notes architectural improvements of roughly 3× less attention compute and a 4.4× smaller KV cache than GLM-5.3. It also flags that GLM 5.3's ow

302.AI

CoverageBenchmark

On August 20, 2026, a listing called "Ox Alpha" appeared on OpenRouter and OpenCode with no owner listed, no model card, a 1M-token context window, and text, image, and video input. Over six days it became the most-used model on the platform, processing about 23 trillion tokens, roughly 2.3 times the volume of the next Weights landed on Hugging Face under an MIT license on August 26, 2026, and Zhipu's Hong Kong shares closed more than 12% higher the next day, roughly 10x the January IPO price. Stripe had agreed to acquire OpenRouter the day before the model appeared, and Stripe CEO Patrick Collison called the stealth listing "very im

302.AI

CoverageBenchmark

GLM 5.3 Flash launched on August 26, 2026 with MIT-licensed weights on Hugging Face, native image and video input, and a 1M-token advertised context window evaluated at 300K, priced at $0.15 in / $0.50 out per million tokens on the Z.ai API. By contrast, the GLM 5.3 flagship shipped on August 14, 2026 as an API-only 74 The comparison highlights that the two launch tables mostly use different benchmark versions, making most cross-model comparisons invalid: the flagship reports Terminal-Bench 3.0 (28.3) while Flash reports Terminal-Bench 2.1 (84.3). DeepSWE v1.1 is the only shared benchmark, where Flash scores 63.4 versus the flagship'

302.AI

Coverage

Z.ai released GLM-5.3-Flash on August 26, 2026, as an open-weight model under the MIT license with 320 billion total parameters and only about 18 billion active parameters per token, designed around the idea of frontier-level capabilities at lower inference cost. The article confirms the model offers a 1-million-token The Medium explainer positions GLM-5.3-Flash as part of a broader shift in the AI model race away from raw parameter counts and toward how much intelligence a model can deliver per unit of compute actually used. It describes Z.ai, formerly known as Zhipu AI, as the creator and frames the model as a frontier-tier sparse

302.AI

CoverageBenchmark

GLM 5.3 Flash is positioned as dramatically cheaper than the GLM 5.3 flagship while delivering close performance: 320B total / 18B active parameters versus 753B / 40B, both with 1M context windows, and API pricing of $0.15 / $0.50 per million tokens for Flash versus $1.40 / $4.40 for the flagship. On Artificial Analysi Contrary to the "Flash = faster" assumption, Artificial Analysis measures output speed at roughly 85 tokens/second for GLM 5.3 versus about 49 tok/s for Flash, with time-to-first-token of 1.46s for Flash versus 1.57s. Flash wins on image understanding, visual coding, browser agents, computer-use agents, and Coding Plan

302.AI

CoverageBenchmark

LLM-Stats ranks GLM-5.3-Flash 17 overall on its composite LLM Stats Score of 50.3 at a blended price of $0.17 per million tokens, placing it between DeepSeek-V4-Flash-0731 and Muse Spark 1.3 on the cost-efficiency frontier. Capability tier breakdowns show Tool Calling at rank 7 of 193 (Top 10%), Reasoning at 16 of 362, Per-benchmark scores sourced from z.ai include GDPval-AA Elo of 1773/3000 (rank 2, Artificial Analysis methodology), CharXiv-R of 0.89/1 (rank 8, with temp=1.0 top_p=0.95 at 256K context), and Terminal-Bench 2.1 results. Performance by conversation depth holds steady around 14.3–14.6 across turns 1 through 2-10 within

302.AI

CoverageBenchmark

BenchLM aggregates verified benchmark coverage for GLM-5.3-Flash with a composite capability score of 66.1/100, a Multimodal & Grounded rank of 10 (noted as "particularly strong for screenshots, documents, charts, and grounded multimodal workflows"), an Agentic rank of 19 of 154 (60.2, 88th percentile, 6 verified bench The aggregator notes 19 of the tracked benchmark slots are empty, and several categories are not measured: Reasoning, Math, Multilingual, and Instruction Following all lack published numbers. Speed and time-to-first-token are not measured, and no comparable first-party hosted API token rate is published, leaving only a

302.AI

CoverageAnalysis

Z.ai shipped GLM-5.3-Flash on August 26, 2026 as a 320B-parameter mixture-of-experts model that activates only 18B parameters per token, runs natively in FP8, and supports a 1,048,576-token context window. It is described as the first natively multimodal model in the GLM-5 series, with native image and video input. Wei For about a week prior to the public reveal, an anonymous listing called "Ox Alpha" topped usage charts on OpenRouter and OpenCode before Bloomberg and other outlets confirmed it was GLM-5.3-Flash from Z.ai. The piece reports that Flash beats GLM-5.2 across six coding and agentic benchmarks at roughly one-tenth the pri

Videos about GLM-5.3-Flash

More models around GLM-5.3-Flash