Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
OrcaRouter logo

Model details

GLM-5.3-Flash (free)

GLM-5.3-Flash is Z.ai's newest flash-tier release and the first model in the GLM-5 line built for native multimodal understanding, accepting both text and images at input. It follows the same open-release philosophy as earlier Z.ai models, shipping under the permissive MIT license so the weights can be downloaded, inspected, and self-hosted. Z.ai positions the Flash variant as a practical, cost-friendly member of the family aimed at developers who want capable reasoning and coding behavior without paying for the largest sibling in the lineup.

Under the hood, GLM-5.3-Flash uses a sparse Mixture-of-Experts design with roughly 320 billion total parameters but only about 18 billion active per token, so each forward pass behaves more like a mid-sized model while retaining a very large knowledge store. The stack registers a hybrid attention architecture that combines sparse and linear attention layers, described by Z.ai as a first for the GLM series, which is intended to keep long-context processing efficient. The model is trained on a large multimodal corpus spanning tens of trillions of tokens and supports feature areas including reasoning, tool use, code generation, and multilingual interaction, making it well suited for assistants, document and image analysis workflows, and agent-style coding tasks where open weights matter.

OrcaRouterz-ai/glm-5.3-flash-freeglm-flash

Quick Info

Powered by
Provider
OrcaRouter
Model key
z-ai/glm-5.3-flash-free
Release date
Aug 26, 2026
Last updated
Aug 26, 2026
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
128,000 tokens
Context window
1,000,000 tokens

Latest news about GLM-5.3-Flash (free)

OrcaRouter

Official sourceAnnouncement

Z.ai introduced GLM-5.3-Flash on August 26, 2026, the first natively multimodal model in the GLM-5 series, with 320B total parameters and 18B active. It combines sparse and linear attention plus Manifold-Constrained Hyper-Connections to cut long-context serving costs while preserving precision on a 30T-token multimodal pre-training corpus. Benchmarks show GLM-5.3-Flash scoring 57 on Artificial Analysis Intelligence Index v4.1.1 at $0.045 per task discounted, outperforming GLM-5.2 by wide margins on DeepSWE v1.1 (63.4 vs 46.2) and AutomationBench (48.8 vs 26.2), and approaching Claude Opus 4.8 on Z.ai Code Bench v1.0. Anonymous testing as ox-alpha on OpenCode and OpenRouter made it the most popular model of the week.

OrcaRouter

Coverage

AI Profit Boardroom details ways to run GLM-5.3-Flash free, including downloading the MIT-licensed open weights or riding the launch promo that halves the already low API price. Z.ai released the model on 26 August 2026 as a 320B-total / 18B-active MoE with a one-million-token context window under the MIT license, nati The page notes the listed context window at 1,310,720 tokens (a figure higher than the 1,048,576 cited elsewhere) with up to 131,072 completion tokens, and positions Flash at roughly a tenth of the flagship GLM-5.3 API rate, or a twentieth during the limited-time discount. Pricing/promo and open-weights availability ma

OrcaRouter

CoverageBenchmark

AI Tech Ledger's review frames GLM-5.3-Flash as combining open weights, multimodal input, an approximately one-million-token context window, and unusually low API pricing — making it relevant for developers building coding agents, long-context workflows, and cost-sensitive AI applications. The article compares Z.ai's o Specific benchmark and pricing details in the supplied excerpt are limited to the introductory framing and cookie-banner text, so the piece's value here is its developer-focused review angle and explicit comparison context. The page identifies GLM-5.3-Flash's open-weights, multimodal, long-context, low-cost combination

OrcaRouter

Coverage

LM Studio added GLM-5.3-Flash to its Bionic agent platform on August 26, 2026, hours after the official unveiling, offering multimodal image and text input with a 1M-token context window. The model is served from US-based servers with zero-data-retention enabled by default and is priced up to 10 times lower than GLM-5.2. GLM-5.3-Flash is a 320B-parameter Mixture-of-Experts model with 18B active parameters, supporting image and text inputs and scoring ahead of GLM-5.2 across Z.ai's highlighted benchmarks while falling within the range of frontier models from Anthropic, OpenAI, Google, and DeepSeek.

OrcaRouter

Coverage

GLM-5.3-Flash, released by Z.ai on August 26, 2026, is a 320-billion-parameter mixture-of-experts model that activates 18 billion parameters per token, with a hybrid sparse and linear attention architecture that Z.ai reports cuts attention compute 3.01x and KV cache size 4.44x versus flagship GLM-5.3. It was pre-traine Local execution as of August 26, 2026 requires a 128GB unified-memory machine and works only at the smallest dynamic quantizations—an estimated 75–80GB at 1-bit and 105–115GB at 2-bit—with the 2-bit path dependent on a Linux memory unlock on Strix Halo hardware; comfortable 4-bit quality needs 192GB–256GB and there is

OrcaRouter

CoverageBenchmark

Artificial Analysis independently confirms GLM-5.3-Flash specs: 320B total / 18B active parameters, 1M-token context window, MIT-licensed open weights, text and image input with text output, and reasoning support. Released August 2026 with Hugging Face weights available. The model scores 42 on the Artificial Analysis Intelligence Index (well above the 18 median for its class) at a $0.25 per-task cost, while running at 48.6 output tokens per second and emitting 180M tokens per Intelligence Index run. Pricing sits at $0.15 per 1M input and $0.50 per 1M output tokens, placing GLM-5.3-Flash among the stronger open-weight models in its size range.

OrcaRouter

CoverageBenchmark

According to Eden AI's benchmark comparison, GLM-5.3-Flash from Z.ai leads its comparison set on Z.ai-reported agentic and vision-heavy benchmarks including Toolathlon Verified (78.4), GDPval-AA v2 Elo (1773), OfficeQA Pro (62.4), and Chartography with Tools (78.0), while not leading Terminal-Bench 2.1, DeepSWE v1.1, o Artificial Analysis gives GLM-5.3-Flash an Intelligence Index of 57, placing it below the separate flagship GLM-5.3 at 60 and below reference models such as Claude Opus 5 (63), Claude Fable 5 (62), GPT-5.6 Sol (61), and Grok 4.6 (61). All comparison figures are Z.ai-reported, and the article frames Flash as strongest i

OrcaRouter

CoverageBenchmark

QCode.cc's guide confirms Z.ai shipped GLM-5.3-Flash on 2026-08-26 as a 320B-total / 18B-active MoE with hybrid sparse and linear attention, accepting text, image, and video input, producing text with tool calling and visual coding support. The official context window is 1,048,576 tokens with 131,072 max completion, wh Launch pricing is $0.15 in / $0.50 out per million tokens (list), with a 50% promo at $0.075 / $0.25 and cache reads at $0.03 ($0.015 during promo), the half-price window closing 2026-09-09 24:00 (UTC+8). Artificial Analysis gives Flash an Intelligence Index of 57 — same tier as Claude Opus 4.8, but below the Fable 5 t

Videos about GLM-5.3-Flash (free)

More models around GLM-5.3-Flash (free)