Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
LLM Gateway logo

Model details

GLM-5.3 Flash (NovitaAI)

GLM-5.3 Flash, released by Z.ai on August 26, 2026, is the first natively multimodal member of the GLM-5 family and was previewed anonymously on OpenRouter as "Ox Alpha" before its public identification. It is built on a newly trained 320-billion-parameter Mixture-of-Experts base that activates roughly 18 billion parameters per token, combining sparse and linear attention layers with Manifold-Constrained Hyper-Connections to keep long-context serving economical. Training drew on a 30-trillion-token multimodal corpus and the weights ship under an MIT license on Hugging Face, with a public vLLM recipe available for local deployment. A native vision encoder handles images and video through the same token stream as text, so the model can ground code, tool calls, and reasoning in visual evidence without a separate vision pipeline.

GLM-5.3 Flash is positioned as an efficient coding and agent workhorse rather than a trimmed flagship, and independent testing confirms the framing. Artificial Analysis places it at the high end of speed among similarly sized open-weight models while noting that it can be verbose on long reasoning tasks, and Z.ai's published results include 84.3 on Terminal-Bench 2.1, 63.4 on DeepSWE v1.1, and 48.8 on AutomationBench, with an independent GDPval-AA v2 Elo that leads several closed alternatives. The combination of a million-token context, tool calling, and native image and video understanding makes it well suited to repository-scale coding agents, document and spreadsheet workflows, and multimodal inspection pipelines where routing routine steps to a lower-cost model preserves budget for the hardest reasoning calls.

LLM Gatewaynovita/glm-5.3-flashglm-flash

Quick Info

Powered by
Provider
LLM Gateway
Model key
novita/glm-5.3-flash
Release date
Aug 26, 2026
Last updated
Aug 26, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.15
Output token cost
$0.50

Limits

Output tokens
131,072 tokens
Context window
1,048,576 tokens

Transparent token rates

Compare GLM-5.3 Flash (NovitaAI) pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GLM-5.3 Flash (NovitaAI)

LLM Gateway

CoverageBenchmark

Z.ai released GLM-5.3-Flash on August 26, 2026, confirming the model previously seen anonymously as "Ox Alpha" on OpenRouter and OpenCode. The release reveals a 320B-parameter mixture-of-experts architecture with only 18B active parameters per token, a 1,048,576-token context window, and native text, image, and video i Z.ai's launch figures include 84.3 on Terminal-Bench 2.1, 63.4 on DeepSWE v1.1, 48.8 on AutomationBench, 78.4 on Toolathlon Verified, and 55.3 on Humanity's Last Exam with tools, with Artificial Analysis independently measuring 57 on its Intelligence Index at roughly 50 tokens per second output. List pricing is $0.15 p

LLM Gateway

Coverage

Z.ai officially revealed GLM-5.3-Flash on August 26, 2026 at 7:42 PM, confirming the model previously previewed under the stealth codename Ox Alpha. Weights were published on Hugging Face under an MIT license, and API documentation went live at docs.z.ai/guides/llm/glm-5.3-flash. The launch post hit 740K+ views within an hour. The model is described as native multimodal with a hybrid sparse and linear attention architecture aimed at efficient coding and long-horizon agent tasks. Open weights enabled third-party work, including an Unsloth Dynamic 3-bit GGUF sized for 128GB of RAM and an OrcaRouter build with refusal-removal edits baked directly into the native block-FP8 weight shards. Z.ai also positioned the model as running entirely on Chinese AI chips.

LLM Gateway

CoverageAnalysis

Z.ai shipped GLM-5.3-Flash on August 26, 2026 as a 320B-parameter mixture-of-experts model that activates only 18B per token, runs natively in FP8, and supports a 1,048,576-token context window. It is the first natively multimodal entry in the GLM-5 series and was previewed anonymously for twelve days on OpenRouter and GLM-5.3-Flash beats GLM-5.2 across six coding and agentic benchmarks at roughly one-tenth the price and approaches Claude Opus 4.8 on long-horizon agent tasks. Weights ship under MIT on Hugging Face, the API is $0.15/$0.50 per million tokens, and the vLLM recipe is already public. Architectural highlights include hybri

LLM Gateway

Coverage

Z.ai formally introduced GLM-5.3-Flash as the first natively multimodal model in the GLM-5 series, combining sparse and linear attention in a hybrid architecture with Manifold-Constrained Hyper-Connections (mHC) and trained on a 30T-token multimodal corpus. The model has 320B total parameters with 18B active and is rel On the Z.ai Code Bench, GLM-5.3-Flash outperforms GLM-5.2 at every effort level and performs on par with Claude Opus 4.8 on coding and agentic workloads. Standard API pricing is $0.15 input, $0.50 output, and $0.03 cached per million tokens, and Artificial Analysis Intelligence Index v4.1.1 scores it 57 at a $0.045-per

LLM Gateway

CoverageBenchmark

Z.ai's GLM-5.3-Flash (open weights, MIT license, released August 2026) scores 42 on the Artificial Analysis Intelligence Index, placing it well above the median for comparable open-weight models. The model uses a 320B total / 18B active-parameter architecture with a 1M-token context window and supports text-and-image input with text output. List pricing is $0.15/$0.50 per million input/output tokens with an 83% cache discount. Output throughput averages roughly 48.6 tokens per second, which Artificial Analysis flags as notably slow compared to peers, and the model is described as very verbose, generating about 180M tokens across the Intelligence Index eval. A reasoning-enabled variant is documented, and the model is suited to efficient coding and long-horizon agent workloads under its hybrid sparse-plus-linear attention design.

LLM Gateway

Coverage

Z.ai introduced GLM-5.3-Flash on August 26, 2026 as a cost-effective frontier model with 320 billion total parameters, 18 billion activated parameters, and a 1 million token context window. The launch was paired with a two-week API discount that halved standard rates for input and output tokens, valid until September 9 Prior to the formal announcement, the model had accumulated significant usage on third-party platforms while operating without attribution under the name "Ox Alpha," which Z.ai later confirmed was a deliberate feedback-gathering step. During this stealth phase, traffic ran entirely on Chinese AI chips, demonstrating ec

LLM Gateway

Coverage

Zhipu AI officially launched GLM-5.3-Flash on August 26, 2026 as a multimodal AI model released under the permissive MIT license, positioning it as a developer-friendly alternative in the vision-language model landscape. The Beijing-based lab made the model immediately available for commercial and research use, removin The Flash variant emphasizes speed and accessibility while maintaining multimodal text-and-image processing capabilities, with a Flash-optimized inference architecture targeting reduced latency. Key technical characteristics include simultaneous text and image input support, integration compatibility with standard tran

LLM Gateway

Coverage

Zhipu AI confirmed on August 26, 2026 that the previously anonymous "Ox Alpha" model was GLM-5.3-Flash, a 320-billion-total / 18-billion-active (320B-A18B) mixture-of-experts release and the first natively multimodal model in the GLM-5 series. The model was priced at roughly a tenth of the flagship GLM-5.3 API rate, wi GLM-5.3-Flash is positioned as one of the cheapest frontier-grade models on any rate card at launch, beating its predecessor GLM-5.2 across benchmarks. Unlike GLM-5.3, whose open weights were still pending at the Flash's launch, GLM-5.3-Flash was self-hostable immediately, marking a notable shift in Zhipu's release str

Videos about GLM-5.3 Flash (NovitaAI)

More models around GLM-5.3 Flash (NovitaAI)