Sulat.com
AI models
Neon logo

Model details

GLM-5.3 Flash

GLM-5.3-Flash is a newly trained Z.ai model rather than a trimmed version of the larger GLM-5.3. Its 320B-parameter mixture-of-experts architecture activates 18B parameters per token and combines linear KDA attention with sparse NoPE MLA layers. Native FP8 weights, an integrated vision encoder, and a 1,048,576-token context are designed to make multimodal and long-context workloads more practical. The model was also trained on a 30-trillion-token multimodal corpus, giving it a separate foundation from the larger sibling while retaining the GLM-5 lineage.

The model is especially well suited to coding, software-engineering agents, terminal automation, and visual tasks such as inspecting rendered interfaces. Its reported results include 84.3 on Terminal-Bench 2.1, 63.4 on DeepSWE v1.1, and 48.8 on AutomationBench v1.0.6, showing strong long-horizon and tool-driven task performance without claiming leadership across every evaluation. The combination of tool use, structured output, native image and video understanding, and a very large context makes it a practical choice for repository-scale analysis and agentic workflows, although the substantial checkpoint and always-on reasoning should be considered when choosing between hosted use and self-hosting.

Neonglm-5-3-flashglm

Quick Info

Powered by
Provider
Neon
Model key
glm-5-3-flash
Release date
Aug 26, 2026
Last updated
Aug 26, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.15
Output token cost
$0.50

Limits

Output tokens
131,072 tokens
Context window
1,048,576 tokens

Transparent token rates

Compare GLM-5.3 Flash pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GLM-5.3 Flash

Neon

CoverageBenchmark

DataCamp reports that Z.ai’s GLM-5.3-Flash uses a 320B-A18B mixture-of-experts design, with 320 billion aggregate parameters and 18 billion active per token, and supports a one-million-token context window. The article describes the exact GLM-5.3-Flash variant as natively multimodal and cost-optimized relative to the l The article gives a Terminal-Bench 2.1 score of 84.3 for GLM-5.3-Flash, compared with 85.0 for Claude Opus 4.8 and 87.4 for GPT-5.6 Terra. It also reports a roughly $0.10 blended price per million tokens, about one-tenth of the cited GLM-5.3 blended price, while noting a 306 GiB FP8 checkpoint that makes lightweight lo

Neon

CoverageAnalysis

A technical deep dive reports that Z.ai released GLM-5.3-Flash on August 26, 2026, as a 320-billion-parameter mixture-of-experts model that activates 18 billion parameters per token. It says the model uses hybrid attention, runs natively in FP8, and supports a 1,048,576-token context window. The same report identifies GLM-5.3-Flash as the first natively multimodal member of the GLM-5 series and says its weights are available on Hugging Face under the MIT license. It also reports a public vLLM recipe, API pricing of $0.15 per million input tokens and $0.50 per million output tokens, and a pre-release anonym

Videos about GLM-5.3 Flash

More models around GLM-5.3 Flash