Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Nvidia logo

Model details

GLM-5.3-Flash

The model overview is temporarily unavailable.

Nvidiaz-ai/glm-5.3-flashglm-flash

Quick Info

Powered by
Provider
Nvidia
Model key
z-ai/glm-5.3-flash
Release date
Aug 26, 2026
Last updated
Aug 26, 2026
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
131,072 tokens
Context window
1,000,000 tokens

Latest news about GLM-5.3-Flash

TokenGo

CoverageBenchmark

A Fello AI technical piece dated September 1, 2026, documents Z.ai's August 26, 2026 release of GLM-5.3-Flash as a 320B-parameter MoE that activates only 18B parameters per token, built on a newly trained base rather than post-trained on GLM-5.2. It is the first natively multimodal model in the GLM-5 series with a 1M-t The piece also reports the API pricing of $0.15 per million input tokens and $0.50 per million output tokens, versus $1.40 and $4.40 for the larger GLM-5.3, and notes that the flagship GLM-5.3's promised 744B open weights are still missing from Z.ai's Hugging Face organization. The key takeaway that "GLM 5.3 Flash is n

TokenGo

CoverageBenchmark

An Ampere comparison piece dated August 27, 2026, places GLM-5.3-Flash side-by-side with the flagship GLM-5.3 using Artificial Analysis independent testing. GLM-5.3 scores 60 on the Intelligence Index against 57 for GLM-5.3-Flash, with GLM-5.3 activating 40B parameters at 753B total versus Flash's 18B at 320B. Both sha On pricing, Flash is roughly one-ninth the cost at $0.15/M input and $0.50/M output versus $1.40/M and $4.40/M for GLM-5.3, and the GLM Coding Plan offers 3× usable quota on Flash versus the 1× reference for the flagship. The piece's category-by-category verdict awards Flash wins on cost-performance, image understandin

OpenRouter

CoverageBenchmark

AI Release Tracker lists GLM-5.3-Flash as an open-weight release from Z.ai on August 26, 2026, arriving twelve days after GLM-5.3, with 320B parameters and a 1M-token context window. The model is described as natively multimodal and designed for agentic coding and tool use, reflecting Z.ai's positioning of the Flash va Headline benchmark scores from the tracker include NL2Repo-Bench 56.3 (described as the best published NL2Repo score among tracked models), DeepSWE 1.1 at 63.4, Terminal-Bench 2.1 at 84.3, Toolathlon-Verified at 78.4, AutomationBench at 48.8, Agent's Last Exam at 26.3 pass@1, and Humanity's Last Exam at 55.3 with tools

TokenGo

CoverageAnalysis

The Local AI Zone technical deep dive, dated August 26, 2026, is the most technically dense of the supplied candidates and corroborates GLM-5.3-Flash as a 320B-total/18B-active mixture-of-experts model that runs natively in FP8 with a 1,048,576-token context window. It is the first natively multimodal model in the GLM- The piece quotes Z.ai's model card stating the model "starts from a newly trained base model, with its architecture and training recipe redesigned around capability and efficiency," explicitly distinguishing it from a post-train of GLM-5.2's 744B base. It also reveals that for about a week prior to launch the model was

Videos about GLM-5.3-Flash

More models around GLM-5.3-Flash