Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Tinfoil logo

Model details

GLM-5.3-Flash

GLM-5.3-Flash, also referred to by community contributors as Ox Alpha, surfaced publicly when its weights were posted on the NVIDIA Developer Forums under the DGX Spark / GB10 Projects board on August 26, 2026. That forum thread, which accumulated thousands of views and active discussion in its first days, signals that the model is being treated by hobbyists and developers as something they can run and inspect themselves rather than a sealed API-only system. An independent Substack guide published two days later framed the model as multimodal and emphasized that it could be hosted locally on consumer hardware, reinforcing the picture of a model designed for hands-on experimentation as much as for hosted use.

Early reporting positioned GLM-5.3-Flash as a multimodal release aimed at bridging open-weight accessibility with frontier-class capability, with third-party commentary highlighting its ability to handle text alongside image and video inputs and to operate on a single workstation rather than requiring a data-center GPU pod. The combination of a million-token context window reported by independent observers and the lightweight "Flash" branding points to an emphasis on long-context throughput and responsive inference, making it a practical fit for builders who want frontier-style multimodal reasoning without committing to a closed provider or specialized accelerators. For teams already comfortable running open models locally, it offers a path to multimodal workloads that can be deployed, fine-tuned, and audited on their own infrastructure.

Tinfoilglm-5-3-flashglm-flash

Quick Info

Powered by
Provider
Tinfoil
Model key
glm-5-3-flash
Release date
Aug 26, 2026
Last updated
Aug 26, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.40
Output token cost
$1.25

Limits

Output tokens
131,072 tokens
Context window
1,048,576 tokens

Transparent token rates

Compare GLM-5.3-Flash pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GLM-5.3-Flash

Tinfoil

CoverageBenchmark

Fello AI's long-form coverage, updated September 1, 2026, confirms the Z.ai announcement that GLM 5.3 Flash shipped on August 26, 2026 with MIT-licensed weights on Hugging Face on day one, a 320B-parameter MoE that activates 18B parameters per token, native multimodality, and a 1M-token context window. Pricing is resta The article also reprints Z.ai's own announcement lines — "Introducing GLM-5.3-Flash — Leading capabilities at a highly competitive price — Natively multimodal with a 1M-token context window — A 320B-A18B model released under the MIT License — Previously previewed as Ox Alpha, running entirely on Chinese AI chips" — an

Tinfoil

CoverageBenchmark

Ampere.sh's GLM 5.3 vs GLM 5.3 Flash comparison, published August 27, 2026, reports independent Artificial Analysis numbers placing GLM 5.3 Flash at 57 on the Intelligence Index against 60 for the flagship GLM 5.3, with output speed of roughly 49 tokens per second for Flash versus about 85 for the full model. It lists The same comparison notes that GLM 5.3 Flash is not actually faster at token generation than GLM 5.3, with time-to-first-token of 1.46s for Flash versus 1.57s for the full model, and awards Flash the wins on cost-performance, image understanding, visual coding, browser agents, computer-use agents, and Coding Plan quota

Tinfoil

CoverageAnalysis

Z.ai released GLM-5.3-Flash on August 26, 2026, and the Local AI Zone technical deep dive documents it as a 320B-total/18B-active mixture-of-experts model that runs natively in FP8 and supports a 1,048,576-token context window. The piece describes it as the first natively multimodal model in the GLM-5 series, built on The same write-up ties the launch to the reveal that GLM-5.3-Flash is "Ox Alpha," the anonymous model that had been quietly serving traffic on OpenRouter and OpenCode for about a week before the announcement. It claims GLM-5.3-Flash beats GLM-5.2 across six coding and agentic benchmarks at roughly one-tenth the price a

Videos about GLM-5.3-Flash

More models around GLM-5.3-Flash