Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Neuralwatt logo

Model details

GLM-5.3 Flash

The model overview is temporarily unavailable.

Neuralwattglm-5.3-flashglm-flash

Quick Info

Powered by
Provider
Neuralwatt
Model key
glm-5.3-flash
Release date
Aug 26, 2026
Last updated
Aug 26, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.15
Output token cost
$0.50

Limits

Output tokens
1,048,560 tokens
Context window
1,048,560 tokens

Transparent token rates

Compare GLM-5.3 Flash pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GLM-5.3 Flash

Neuralwatt

CoverageBenchmark

Z.ai released GLM-5.3-Flash on August 26, 2026, with MIT-licensed weights on Hugging Face on day one, and the article frames it as the first natively multimodal model in the GLM-5 series built on a newly trained base rather than a post-train of GLM-5.2. It is a 320B-total / 18B-active mixture of experts with a 1M-token Pricing is $0.15 per million input tokens and $0.50 output, compared with $1.40 and $4.40 for the GLM-5.3 flagship. The piece also documents the Ox Alpha reveal—the anonymous model that had been serving traffic on OpenRouter with a 1M-token context window—and notes that the flagship 744B GLM-5.3 weights promised at lau

Neuralwatt

CoverageBenchmark

Z.ai's GLM-5.3-Flash (the model Zhipu AI ships internationally as Z.ai) was confirmed on August 26, 2026 as the identity of the previously anonymous "Ox Alpha" model that had topped OpenRouter usage, processing roughly 23 trillion tokens in six days—about 2.3x the volume of the next model—and ending DeepSeek's 56-day r Weights landed on Hugging Face under an MIT license the same evening, and Zhipu's Hong Kong shares closed more than 12% higher the next day at roughly 10x January's IPO price. The piece situates the release against the OpenRouter acquisition context (Stripe agreed to buy OpenRouter the day before "Ox Alpha" appeared) a

Neuralwatt

Coverage

Z.ai released GLM-5.3-Flash on August 26, 2026 as a 320B-total / 18B-active mixture-of-experts model with native multimodal input, a one-million-token context window, and an MIT-licensed open-weight release. The article lists the OpenRouter model's context window as 1,310,720 tokens with up to 131,072 completion tokens The piece reconstructs the Ox Alpha reveal: an anonymous free model appeared on a third-party API platform on August 20, 2026 with text/image/video input and a million-token context, the community fingerprinted it back to the GLM family within 48 hours, and Z.ai's August 26 announcement confirmed the preview was theirs

Neuralwatt

Coverage

Z.ai released GLM-5.3-Flash on August 26, 2026 as an open-weight model under MIT license, with 320 billion total parameters, about 18 billion active per token, multimodal inputs (text, image, video), and a context window of up to one million tokens. The piece frames it as delivering frontier-level capabilities without The article traces the "Ox Alpha" origin story: before Z.ai officially announced the model, an anonymous preview labeled Ox Alpha had been quietly served through OpenRouter and OpenCode to gather real-world feedback, and Z.ai later confirmed the identity on August 26, 2026. It positions GLM-5.3-Flash as a new member of

Neuralwatt

CoverageBenchmark

An independent Artificial Analysis comparison scores GLM-5.3 at 60 on the Intelligence Index versus 57 for GLM-5.3-Flash, while Flash activates only 18B parameters against GLM-5.3's 40B and costs roughly one-ninth as much at normal API pricing. Specs listed: GLM-5.3 is 753B total / 40B active at ~85 tok/s with a 1.57s The article flags that GLM-5.3-Flash is the only one of the two with native image input and native multimodal support, and is the one developers can actually download today—GLM-5.3's public weights are listed as "coming soon." It also notes the "GLM-5.3-Flash" name is a misnomer in that it is not a trim of GLM-5.2's ba

Neuralwatt

CoverageAnalysis

Z.ai shipped GLM-5.3-Flash on August 26, 2026, as a 320-billion-parameter mixture-of-experts model that activates only 18B parameters per token, runs natively in FP8, and supports a 1,048,576-token context window. The model is described on the card as the first natively multimodal model in the GLM-5 series, and its arc Weights were released under MIT license on Hugging Face the same day, with the vLLM recipe already public, and Z.ai priced the API at $0.15 per million input tokens and $0.50 output—a fraction of GLM-5.3's $1.40/$4.40 rates. According to the article, GLM-5.3-Flash beats GLM-5.2 across six coding and agentic benchmarks

Videos about GLM-5.3 Flash

More models around GLM-5.3 Flash