Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Volcengine Ark logo

Model details

GLM-5.3-Flash

The model overview is temporarily unavailable.

Volcengine Arkglm-5-3-flash-260828glm-flash

Quick Info

Powered by
Provider
Volcengine Ark
Model key
glm-5-3-flash-260828
Release date
Aug 26, 2026
Last updated
Aug 26, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.11875
Output token cost
$0.41563

Limits

Output tokens
131,072 tokens
Context window
1,000,000 tokens

Transparent token rates

Compare GLM-5.3-Flash pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GLM-5.3-Flash

Volcengine Ark

CoverageBenchmark

Z.ai released GLM-5.3-Flash on August 26, 2026, shipping MIT-licensed weights on Hugging Face the same day. It is a 320B-total / 18B-active mixture-of-experts built on a newly trained base rather than a post-train of GLM-5.2, and it is the first natively multimodal model in the GLM-5 series with a 1M-token context and API pricing is $0.15 per million input tokens and $0.50 per million output tokens, roughly one-tenth the flagship GLM-5.3 rate of $1.40 input and $4.40 output. The article notes that the promised open-weight release of GLM-5.3's 744B weights has not yet materialized, making GLM-5.3-Flash what developers actually have a

Volcengine Ark

Coverage

Z.ai released GLM-5.3-Flash on August 26, 2026, as a 320-billion-parameter model with 18 billion active, natively multimodal, with a one-million-token context window under the MIT license. The OpenRouter model listing shown in the article reports a 1,310,720-token context window with up to 131,072 completion tokens and The article chronicles the Ox Alpha origin story: an anonymous model appeared on a third-party API platform on August 20 with a million-token context and free pricing, was fingerprinted back to the GLM family within 48 hours, and was confirmed by Z.ai's August 26 announcement as GLM-5.3-Flash. The piece positions the m

Volcengine Ark

Coverage

Z.ai released GLM-5.3-Flash on August 26, 2026, as the newest member of the GLM-5 family, explicitly framed around delivering frontier-level capabilities without frontier-level inference cost. The model is a 320-billion-parameter design that activates only about 18 billion parameters, supports multimodal inputs, and of According to the article, GLM-5.3-Flash sits at the intersection of the increasing competition over intelligence-per-compute and the shift toward open-weight frontier-class models. The piece emphasizes that the release concentrates capability into a small active-parameter footprint, which is expected to materially redu

Volcengine Ark

CoverageBenchmark

Drawing on Artificial Analysis independent testing, GLM-5.3 scores 60 on the Intelligence Index compared with 57 for GLM-5.3-Flash, yet GLM-5.3-Flash costs roughly one-ninth as much at normal API pricing and activates only 18B parameters versus GLM-5.3's 40B. Both models share a 1M-token context, support reasoning and The comparison places GLM-5.3 as the pick for maximum coding quality and generation speed, and GLM-5.3-Flash as the better value for high-volume applications, multimodal agents, and cost-performance-optimized production workloads, with a 3× usable Coding Plan quota advantage. Open weights are currently available only f

Volcengine Ark

CoverageBenchmark

GLM-5.3-Flash ranks 17 on the LLM Stats composite score and places Top 10% in Tool Calling (#7 of 193), Top half in Reasoning (#16 of 362), Chat (#21 of 117), Vision (#24 of 208), Coding (#32 of 266), and Math (#38 of 327). Per-dataset scores include GDPval-AA 1773/3000 from Artificial Analysis, CharXiv-R 0.89 sourced On the cost-efficiency chart, GLM-5.3-Flash is listed at a blended $0.17 per million tokens with an LLM Stats Score of 50.3, positioning it favorably against both budget models like Gemma 4 E4B ($0.024 / 13.8) and premium options like GPT-6 Astra ($11.9 / 59.9). Performance by conversation depth shows stable scores acr

Volcengine Ark

CoverageAnalysis

Z.ai shipped GLM-5.3-Flash on August 26, 2026, as a 320B-parameter mixture-of-experts that activates 18B per token, runs natively in FP8, supports a 1,048,576-token context window, and is the first natively multimodal model in the GLM-5 series. The deep dive confirms the model was anonymously running on OpenRouter as " Benchmark coverage shows GLM-5.3-Flash beating GLM-5.2 across six coding and agentic benchmarks at roughly one-tenth the price while approaching Claude Opus 4.8 on long-horizon agent tasks. The architecture uses a hybrid sparse-and-linear attention design intended for efficiency, with explicit Z.ai creator attribution

Videos about GLM-5.3-Flash

More models around GLM-5.3-Flash