Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Inco logo

Model details

GLM-5.3-Flash

GLM-5.3-Flash is a large multimodal model from Z.ai designed to deliver strong coding and agentic capability at a lower serving cost than typical frontier systems. It carries 320B total parameters with just 18B active, a sparse-plus-linear hybrid attention design intended to preserve long-context precision while cutting inference expense, and a newly trained base model rebuilt around efficiency. Z.ai describes it as the first natively multimodal release in the GLM-5 series, supported by a 30T-token multimodal pre-training corpus and architectural changes such as manifold-constrained hyper-connections that improve scaling efficiency. Open weights are published on Hugging Face under the zai-org name, enabling local deployment and experimentation for developers who want frontier-class behavior without closed-API lock-in.

The model's intended use centers on coding assistants, agentic workflows, and general multimodal reasoning where long context matters. Z.ai positions GLM-5.3-Flash as outperforming its GLM-5.2 predecessor across a range of benchmarks and real-world workloads while approaching Claude Opus 4.8 on coding and agentic evaluations, framing the result as frontier intelligence at flash-tier cost. The model gained traction first under the anonymous alias "Ox Alpha," where it quietly topped usage charts on OpenRouter and OpenCode for about a week and drew attention from developers and executives before Z.ai confirmed authorship. Practical fit includes teams building code generation, tool-driven agents, and document or media understanding pipelines that benefit from an open-weight, long-context model with multimodal input and structured output.

Incoglm-5.3-flash:fastglm-flash

Quick Info

Powered by
Provider
Inco
Model key
glm-5.3-flash:fast
Release date
Aug 26, 2026
Last updated
Aug 26, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.15
Output token cost
$0.50

Limits

Output tokens
131,072 tokens
Context window
1,000,000 tokens

Transparent token rates

Compare GLM-5.3-Flash pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GLM-5.3-Flash

Inco

CoverageBenchmark

Z.ai released GLM-5.3-Flash on August 26, 2026, as a 320-billion-parameter mixture-of-experts model that activates just 18 billion parameters per token, making it the first natively multimodal model in the GLM-5 series and the first built on a new base since GLM-5.2. The model supports a 1M-token context window and use API pricing is set at $0.15 per million input tokens and $0.50 per million output tokens, compared with $1.40 and $4.40 for the flagship GLM-5.3, roughly a one-tenth cost reduction. Notably, the open-weight release of GLM-5.3's promised 744B parameters has not materialized: Z.ai's Hugging Face organisation has no GLM-5

Inco

CoverageBenchmark

In independent Artificial Analysis testing, GLM-5.3 scores 60 on the Intelligence Index compared with 57 for GLM-5.3-Flash, while GLM-5.3-Flash costs roughly one-ninth as much at normal API pricing and activates only 18B parameters versus 40B for GLM-5.3. Counter-intuitively, GLM-5.3-Flash is not faster at token genera The spec comparison confirms both models are from Z.ai with 1M context windows, reasoning and tool use support, and the same multimodal context length, but only GLM-5.3-Flash offers native image input and is natively multimodal. GLM-5.3-Flash offers 3× usable quota in the Coding Plan versus 1× reference for GLM-5.3, an

Inco

CoverageBenchmark

BenchLM's aggregator data ranks GLM-5.3-Flash at capability score 66.1/100 (field median 56.3, placing it 32nd of 230 ranked models), with multimodal ranking at 10th, making it particularly strong for screenshots, documents, charts, and grounded multimodal workflows. Verified benchmark coverage includes Agentic (rank 1 The model supports a 1M-token context window and is tracked across 496 destinations with 446 benchmarks, though Reasoning, Math, Multilingual, and Instruction-Following categories remain unmeasured. Pricing is listed as self-hosted with infrastructure cost varying (input median $0.95), and no comparable first-party hos

Inco

CoverageAnalysis

A technical deep dive confirms GLM-5.3-Flash ships as a 320B-parameter MoE activating 18B per token, running natively in FP8 with a 1,048,576-token context window and hybrid sparse+linear attention. The architecture was trained from a newly designed base model rather than post-trained from GLM-5.2's 744B base, with the Weights are released under MIT on Hugging Face at launch, and the API pricing is $0.15/$0.50 per million tokens. The piece documents the Ox Alpha identity reveal, noting the mystery model had been topping OpenRouter usage charts and drawing praise from figures like Patrick Collison before its August 26 launch. The tech

Videos about GLM-5.3-Flash

More models around GLM-5.3-Flash