Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
LLM Gateway logo

Model details

GLM-5.3 Flash (Inference.net)

GLM-5.3 Flash is a model from the GLM family released by Z.ai, with sources confirming its launch in late August 2026. The Mercatus pricing analysis identifies Z.ai as the creator operating the official API, and the model's name places it within the broader GLM Flash lineage that has historically targeted faster, lower-cost inference compared with flagship-tier siblings. Coverage in independent comparison articles treats GLM-5.3 Flash as a peer to competing open-weight Flash-tier models, suggesting an intended positioning as a speed- and cost-optimized variant rather than a maximum-quality offering.

Practically, GLM-5.3 Flash is aimed at developers and teams who want access to a recent Z.ai open-weight model through flexible hosting, with multiple third-party providers serving the weights and pricing the model below Z.ai's official list rate. This distribution pattern indicates a design that tolerates broad self-hosting and redistribution, fitting use cases such as high-volume text generation, chat applications, and integrations where open-weight licensing is a requirement. Buyers evaluating the model should weigh the speed and cost advantages suggested by its Flash designation against the limited independent verification currently available about its benchmark performance and architectural details.

LLM Gatewayinference.net/glm-5.3-flashglm-flash

Quick Info

Powered by
Provider
LLM Gateway
Model key
inference.net/glm-5.3-flash
Release date
Aug 26, 2026
Last updated
Aug 26, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.09
Output token cost
$0.28

Limits

Output tokens
128,000 tokens
Context window
1,048,576 tokens

Transparent token rates

Compare GLM-5.3 Flash (Inference.net) pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GLM-5.3 Flash (Inference.net)

LLM Gateway

Coverage

Z.ai officially released GLM-5.3-Flash on August 26, 2026, as a newly trained model rather than a post-train of GLM-5.2, trained on a 30T-token multimodal corpus. The architecture combines sparse and linear attention with Manifold-Constrained Hyper-Connections, reducing active parameters from 32B to 18B and cutting attention compute roughly 3.0× versus GLM-5.3 with a 4.4× smaller KV-cache. The model serves efficiently at 1M-token context lengths and was anonymously evaluated as Ox Alpha on Chinese AI accelerators before launch. Z.ai reports GLM-5.3-Flash outperforms GLM-5.2 across coding and agentic benchmarks, scoring 84.3 on Terminal Bench 2.1, 63.4 on DeepSWE v1.1, 56.3 on NL2Repo, 78.4 on Toolathlon Verified, 48.8 on AutomationBench v1.0.6, 26.3 on Agents' Last Exam, 55.3 on HLE with Tools, and 1773 on GDPval-AA v2. The IndexPool mechanism further compresses cached keys at million-token contexts, supporting long-horizon agent workloads. Weights ship under MIT, with vLLM and SGLang recipes available for self-hosting.

LLM Gateway

CoverageBenchmark

BuildFastWithAI's review confirms Z.ai's August 26, 2026 release of GLM-5.3-Flash as a 320B-total/18B-active MoE with text, image, and video input, a 1M-token context window, and MIT-licensed open weights. The review attributes the model to Z.ai and reports list pricing at $0.15 per million input tokens and $0.50 per million output tokens, with a 50% launch promotion running through September 9, 2026. Support spans vLLM, SGLang, and TokenSpeed serving frameworks. The review aggregates Z.ai's reported benchmark results, including 84.3 on Terminal-Bench 2.1, 63.4 on DeepSWE v1.1, 48.8 on AutomationBench, 78.4 on Toolathlon Verified, and 55.3 on Humanity's Last Exam with tools. Artificial Analysis is cited at 57 on its Intelligence Index with roughly 50 output tokens per second at list pricing. The piece positions GLM-5.3-Flash as a value-leading open model for coding and agentic workloads.

LLM Gateway

Coverage

Zhipu AI's shares jumped after the company revealed that the viral anonymous model "Ox Alpha" was its newly launched GLM-5.3-Flash, marking the model's official open-weight debut. The system reportedly ran entirely on 100,000 domestically produced chips during the high-profile stealth trial. The reveal positions GLM-5.3-Flash as evidence that China can scale AI on homegrown semiconductors amid Beijing's effort to cut reliance on US hardware. The model's usage during the anonymous phase set platform traffic records before its public attribution to Zhipu.

LLM Gateway

CoverageAnalysis

Local AI Zone's deep-dive confirms GLM-5.3-Flash shipped August 26, 2026, as a 320B-total/18B-active MoE running natively in FP8 with a 1,048,576-token context window. It is the first natively multimodal member of the GLM-5 series, with weights released under MIT on Hugging Face and a public vLLM recipe. The model had appeared anonymously on OpenRouter as Ox Alpha for roughly twelve days before its formal reveal by Bloomberg. The technical review highlights GLM-5.3-Flash beating GLM-5.2 across six coding and agentic benchmarks at roughly one-tenth the price while approaching Claude Opus 4.8 on long-horizon agent tasks. Notably, the GLM-5.3 flagship's 744B parameters remain unreleased on Z.ai's Hugging Face organization, with the Flash variant shipping first. The architecture is explicitly redesigned rather than distilled, using a fresh base checkpoint.

LLM Gateway

CoverageBenchmark

Artificial Analysis rates GLM-5.3-Flash at 42 on its Intelligence Index, placing it well above the comparable-model median of 14 among open-weight models. The model runs at roughly 44.8 output tokens per second and supports text and image input with text output at a 1M-token context window. Pricing is listed at $0.15 per million input tokens and $0.50 per million output tokens, with an 83% cache discount and $0.25 average cost per Intelligence Index task. Specifications confirmed by Artificial Analysis include 320B total parameters with 18B active per token, MIT-licensed open weights, and a reasoning-enabled variant alongside potential non-reasoning siblings. The page notes high verbosity at 180M output tokens during Intelligence Index evaluation versus a 140M median. GLM-5.3-Flash is classified within the open-weights small-to-medium parameter tier for independent comparison.

LLM Gateway

CoverageBenchmark

Zhipu confirmed that its anonymous "Ox Alpha" model is GLM-5.3-Flash (320B-A18B), which scored 57 on the Artificial Analysis Intelligence Index, tying Claude Opus 4.8 and surpassing DeepSeek V4 Pro's 53. During a stealth trial on OpenRouter and OpenCode, the model accumulated over 50 trillion tokens in five days, with usage reportedly surpassing twice that of DeepSeek on those platforms. The full inference workload was served by more than 100,000 domestic Chinese chips, with Zhipu stating that hardware efficiency and per-token cost reached levels comparable to mainstream NVIDIA GPUs. The disclosure fueled a more than 9% intraday stock surge for Zhipu on August 26, highlighting market focus on the domestic-chip scaling narrative behind GLM-5.3-Flash.

LLM Gateway

Coverage

Zhipu AI has released GLM-5.3-Flash as an open-source model powered entirely by a cluster of 100,000 domestically produced Chinese chips, underscoring China's push for homegrown AI infrastructure. The model features a 300-billion-parameter architecture, supports a 1-million-token context window, and is described as the first native multimodal release in the GLM-5 series. GLM-5.3-Flash scored 57 on the Artificial Analysis Intelligence Index, on par with Anthropic's Claude Opus 4.8. Zhipu priced it at $0.15 per million input tokens and $0.50 per million output tokens, roughly one-tenth of GLM-5.3 and about one-fortieth of Claude Opus 4.8, with cost gains attributed to a hybrid sparse-plus-linear attention architecture that preserves long-context accuracy while reducing serving cost.

Videos about GLM-5.3 Flash (Inference.net)

More models around GLM-5.3 Flash (Inference.net)