Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Umans AI logo

Model details

GLM 5.3 Flash

The model overview is temporarily unavailable.

Umans AIumans-glm-5.3-flashglm-flash

Quick Info

Powered by
Provider
Umans AI
Model key
umans-glm-5.3-flash
Release date
Aug 26, 2026
Last updated
Aug 26, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.15
Output token cost
$0.50

Limits

Output tokens
131,071 tokens
Context window
1,048,576 tokens

Transparent token rates

Compare GLM 5.3 Flash pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GLM 5.3 Flash

Umans AI Coding Plan

CoverageBenchmark

Fello AI's coverage confirms Z.ai's August 26, 2026 release of GLM 5.3 Flash and frames it as a distinct 320B-A18B MoE built on a new base, not a distilled flagship. It is the first natively multimodal model in the GLM-5 series, advertises a 1M-token context, and uses a hybrid sparse-and-linear attention design. MIT-li The article highlights the pricing gap between Flash and the flagship: $0.15 per million input tokens and $0.50 per million output for Flash, versus $1.40 and $4.40 for GLM 5.3. It also notes that GLM 5.3's promised 744B open-weight release has not materialised, and Z.ai's Hugging Face organisation has no GLM 5.3 repos

Umans AI Coding Plan

CoverageBenchmark

Yotta Labs' comparison (August 27, 2026) places GLM 5.3 Flash alongside Z.ai's August 14 flagship and surfaces a key analytical point: the two launch benchmark tables use mostly different test versions, making direct comparison misleading. Terminal-Bench 3.0 is reported for the flagship at 28.3 while Flash reports Term The article documents access asymmetry: GLM 5.3 launched as API-only while GLM 5.3 Flash ships with MIT-licensed open weights, native image and video input, and pricing around one-tenth of the flagship's. Flash is positioned as self-hostable with an 8-GPU node floor. The piece also flags that Flash spent a week running

Umans AI Coding Plan

Coverage

The Medium summary (August 26, 2026) by Greek Ai on CodeToDeploy restates the core GLM-5.3-Flash launch facts: 320 billion total parameters with about 18 billion active, multimodal input support, a 1M-token context window, and open weights under the MIT license, all released by Z.ai on August 26, 2026. The framing posi The piece is largely a synopsis of Z.ai's announcement text rather than an independent technical analysis. It repeats the value proposition of dramatically lower inference costs while delivering competitive performance, but adds limited new benchmark or architectural detail beyond what the launch announcement itself pr

Umans AI Coding Plan

CoverageBenchmark

Ampere's head-to-head (August 27, 2026) cites Artificial Analysis intelligence-index scores of 60 for GLM 5.3 and 57 for GLM 5.3 Flash, alongside output-speed measurements of roughly 85 tokens per second for the flagship versus about 49 tok/s for Flash. The comparison frames Flash as dramatically cheaper and natively m The spec table lists GLM 5.3 as a 753B-total / 40B-active MoE and Flash as 320B-total / 18B-active, both with a 1M context window. Pricing is $1.40/M input and $4.40/M output for the flagship against $0.15/M input and $0.50/M output for Flash, with Flash weights available now and the flagship's weights listed as "comin

Umans AI Coding Plan

CoverageAnalysis

GLM-5.3-Flash was shipped by Z.ai on August 26, 2026, as a 320-billion-parameter mixture-of-experts model that activates only 18B parameters per token and runs natively in FP8, according to a technical deep dive on Local AI Zone. It is the first natively multimodal model in the GLM-5 series, supports a 1,048,576-token The article details the developer-facing surface: weights are released under the MIT license on Hugging Face, the vLLM serving recipe is already public, and the Z.ai API is priced at $0.15 per million input tokens and $0.50 per million output tokens. On coding and agentic benchmarks the model beats GLM-5.2 across six t

Videos about GLM 5.3 Flash

More models around GLM 5.3 Flash