Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Pareto Inference logo

Model details

GLM-5.3-Flash

GLM-5.3-Flash is a new iteration of the GLM series from Z.ai, revealed on August 26, 2026 after a week of community speculation about a stealth-listed model that had appeared on OpenRouter under the identifier stealth/ox-alpha. It is described as the first natively multimodal member of the GLM-5 series, with native support for text, image, and video input. During its preview phase on OpenRouter, the model offered a one-million-token context window and unlimited free access, which fueled rapid adoption and attention before the official identity was disclosed.

Architecturally, GLM-5.3-Flash uses a Mixture-of-Experts design with 320 billion total parameters and 18 billion active parameters per token, a sparse activation pattern that aims to balance large-model capacity with more efficient inference. Early reactions from developers running it on DGX Spark hardware characterized it favorably relative to its predecessor, GLM-5.2, suggesting practical quality gains alongside the expansion into image and video understanding. The combination of multimodal native input, a large context window, and an MoE efficiency profile positions it as a versatile general-purpose model suited to long-context document analysis, vision-and-language tasks, and mixed-modality workflows.

Pareto Inferencez-ai/glm-5.3-flashglm-flash

Quick Info

Powered by
Provider
Pareto Inference
Model key
z-ai/glm-5.3-flash
Release date
Aug 26, 2026
Last updated
Aug 26, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.03
Output token cost
$0.10

Limits

Output tokens
131,072 tokens
Context window
1,000,000 tokens

Transparent token rates

Compare GLM-5.3-Flash pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GLM-5.3-Flash

IteraCompute

Official sourceAnnouncement

Z.ai officially announced GLM-5.3-Flash on August 26, 2026, describing it as the first natively multimodal model in the GLM-5 series. The model is a mixture-of-experts with 320B total parameters and 18B active per token, offering a 1,048,576-token context window and supporting image and video input. Z.ai states GLM-5.3 On the Artificial Analysis Intelligence Index v4.1.1, Z.ai reports GLM-5.3-Flash scores 57 at a discounted cost of $0.045 per task. On coding benchmarks it posts 63.4 on DeepSWE v1.1 (up from 46.2 for GLM-5.2) and 48.8 on AutomationBench (up from 26.2), and on Z.ai Code Bench v1.0 it scores 29.0 at max effort versus 29

TokenGo

CoverageBenchmark

A Fello AI technical piece dated September 1, 2026, documents Z.ai's August 26, 2026 release of GLM-5.3-Flash as a 320B-parameter MoE that activates only 18B parameters per token, built on a newly trained base rather than post-trained on GLM-5.2. It is the first natively multimodal model in the GLM-5 series with a 1M-t The piece also reports the API pricing of $0.15 per million input tokens and $0.50 per million output tokens, versus $1.40 and $4.40 for the larger GLM-5.3, and notes that the flagship GLM-5.3's promised 744B open weights are still missing from Z.ai's Hugging Face organization. The key takeaway that "GLM 5.3 Flash is n

IteraCompute

CoverageBenchmark

Ampere.sh published a third-party comparison of GLM 5.3 versus GLM 5.3 Flash on August 27, 2026, drawing on independent Artificial Analysis testing. GLM 5.3 scores 60 on the Intelligence Index compared with 57 for GLM 5.3 Flash, but Flash costs around one-ninth as much at normal API pricing and activates only 18B param The comparison frames GLM 5.3 Flash as the better value for high-volume applications and multimodal agents, while GLM 5.3 remains preferable where maximum coding quality and generation speed matter more than cost. Specs listed include 753B total / 40B active parameters for GLM 5.3 versus 320B / 18B for Flash, both with

IteraCompute

CoverageRelease Notes

MarkTechPost covered Z.ai's August 26, 2026 release of GLM-5.3-Flash, characterizing it as the first natively multimodal model in the GLM-5 series and Z.ai's cheapest capable coding model to date. The model is described as a 320B-total/18B-active mixture-of-experts with a 1,048,576-token context window and image and vi The article notes GLM-5.3-Flash was first tested anonymously as "Ox Alpha" on OpenCode and OpenRouter, with all serving done on domestically produced Chinese AI chips. Deployment options include the MIT-licensed FP8 weights on Hugging Face, which are roughly 306 GiB before KV cache, and a hosted API that is already pri

IteraCompute

CoverageBenchmark

Kingy AI published an original hands-on review of GLM-5.3-Flash on August 26, 2026, characterizing it as an unusually credible low-cost agent model. The piece confirms the core specification: a 320B-total, 18B-active mixture-of-experts with a claimed 1,048,576-token context, native multimodal capabilities, MIT-licensed Kingy AI ran five original tasks through Z.ai's public chat interface — strict structured output, executable JavaScript, instruction-injection resistance, exact cost arithmetic, and retrieval from a 300-record packet — and reports GLM-5.3-Flash passed all five on the first attempt, with generated JavaScript passing 14

OpenRouter

CoverageBenchmark

AI Release Tracker lists GLM-5.3-Flash as an open-weight release from Z.ai on August 26, 2026, arriving twelve days after GLM-5.3, with 320B parameters and a 1M-token context window. The model is described as natively multimodal and designed for agentic coding and tool use, reflecting Z.ai's positioning of the Flash va Headline benchmark scores from the tracker include NL2Repo-Bench 56.3 (described as the best published NL2Repo score among tracked models), DeepSWE 1.1 at 63.4, Terminal-Bench 2.1 at 84.3, Toolathlon-Verified at 78.4, AutomationBench at 48.8, Agent's Last Exam at 26.3 pass@1, and Humanity's Last Exam at 55.3 with tools

TokenGo

CoverageAnalysis

The Local AI Zone technical deep dive, dated August 26, 2026, is the most technically dense of the supplied candidates and corroborates GLM-5.3-Flash as a 320B-total/18B-active mixture-of-experts model that runs natively in FP8 with a 1,048,576-token context window. It is the first natively multimodal model in the GLM- The piece quotes Z.ai's model card stating the model "starts from a newly trained base model, with its architecture and training recipe redesigned around capability and efficiency," explicitly distinguishing it from a post-train of GLM-5.2's 744B base. It also reveals that for about a week prior to launch the model was

Kilo Gateway

Official sourceDocumentation

Z.AI's official developer documentation confirms GLM-5.3-Flash as the first native multimodal model in the GLM-5 series, delivering stronger intelligence than GLM-5.2 at lower cost. It pairs a 320B total / 18B activated hybrid sparse-plus-linear attention architecture, which cuts attention compute and KV cache by roughly 3x and 4.4x versus GLM-5.3 while preserving long-context quality. The model is now generally available on the GLM Coding Plan with 3x the prior quota. The page documents native visual coding, where the model observes interfaces and rendered results inside the coding loop to coordinate tasks across code, browsers, and GUIs, including frontend, game development, and Blender 3D scenes, while also supporting Office, financial research, and document workflows that produce PPTX, PDF, DOCX, and XLSX outputs. A faster sibling, GLM-5.3-FlashX, reaches 200 tokens/s. The model accepts video, image, text, and file inputs with a 1M-token context and 128K max output, using temperature 1, top_p 0.95, and reasoning effort max for recommended settings.

IteraCompute

CoverageBenchmark

Artificial Analysis provides an independent benchmark page for GLM 5.3 Flash, listing it as an open-weights model from the Z AI lab released in August 2026. The model scores 42 on the Artificial Analysis Intelligence Index, placing it well above the median of 18 among comparable models, and generates output at 89.1 tok Technical specifications confirm 320B total parameters with 18B active per token, a 1M-token context window, support for text and image input with text output, reasoning capability (with a note that a non-reasoning variant may also exist), and an MIT license with weights hosted on Hugging Face. The Artificial Analysis

Videos about GLM-5.3-Flash

More models around GLM-5.3-Flash