Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Bothub logo

Model details

GLM-5.3-Flash

GLM-5.3-Flash is an open-weight release from Z.ai, built as a successor to the GLM-4.5 line and trained from a new base model on a 30T-token multimodal corpus. Before its official reveal, the model circulated anonymously as Ox Alpha on OpenRouter and OpenCode, where it topped usage charts and drew attention from developers and press for its strong coding and reasoning performance under real-world traffic. That community-driven validation gave Z.ai a credible launch narrative once the model's identity was confirmed.

The architecture is the model's defining feature. GLM-5.3-Flash keeps roughly the same total parameter count as the GLM-4.5 series (around 320B) but drops active parameters from 32B to 18B and shrinks the network from 92 to 45 layers, making it far more efficient at inference. To handle long contexts cheaply, it blends linear attention for local dependencies with sparse attention that uses a lightweight indexer to pull relevant global information, and adds Manifold-Constrained Hyper-Connections for cleaner scaling. The result is dramatically reduced attention compute and a smaller KV cache, while still delivering frontier coding and multimodal reasoning that fits local deployment workflows.

Bothubglm-5.3-flashglm-flash

Quick Info

Powered by
Provider
Bothub
Model key
glm-5.3-flash
Release date
Aug 26, 2026
Last updated
Aug 26, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.12
Output token cost
$0.44

Limits

Output tokens
131,072 tokens
Context window
1,000,000 tokens

Transparent token rates

Compare GLM-5.3-Flash pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GLM-5.3-Flash

Bothub

CoverageBenchmark

Z.ai released GLM-5.3-Flash on August 26, 2026, shipping it as a 320-billion-parameter mixture-of-experts model that activates just 18 billion parameters per token, with MIT-licensed weights posted to Hugging Face on day one. According to Z.ai's announcement quoted in the piece, it is the first natively multimodal mode API pricing for GLM-5.3-Flash is set at $0.15 per million input tokens and $0.50 per million output tokens, roughly one-ninth the cost of the GLM-5.3 flagship's $1.40/$4.40 rates. The piece also notes that Z.ai's promised open-weight release of the GLM-5.3 flagship's 744B parameters, expected around launch, has not mat

Bothub

CoverageBenchmark

A side-by-side comparison between GLM-5.3 and GLM-5.3-Flash draws on Artificial Analysis testing and Z.ai's published specs. GLM-5.3 scores 60 on the Intelligence Index versus 57 for GLM-5.3-Flash, but GLM-5.3-Flash activates only 18B parameters (vs. 40B) and costs roughly one-ninth as much at normal API pricing ($0.15 A notable counter-narrative finding: GLM-5.3-Flash is not faster at generation than the flagship. Artificial Analysis measures output at roughly 49 tokens per second for Flash versus 85 tok/s for GLM-5.3, with time-to-first-token nearly identical (1.46s vs. 1.57s). The article concludes GLM-5.3-Flash is the better valu

Bothub

CoverageAnalysis

GLM-5.3-Flash shipped on August 26, 2026 as a 320B-parameter mixture-of-experts model activating 18B parameters per token, trained natively in FP8 with a 1,048,576-token context window and described as the first natively multimodal model in the GLM-5 series. The deep-dive confirms it is built from a newly trained base Benchmark coverage shows GLM-5.3-Flash beating GLM-5.2 across six coding and agentic benchmarks at roughly one-tenth the price, and approaching Claude Opus 4.8 on long-horizon agent tasks. The article surfaces developer-facing details: MIT-licensed weights on Hugging Face, API pricing at $0.15/$0.50 per million tokens,

Videos about GLM-5.3-Flash

More models around GLM-5.3-Flash