Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
LLM Gateway logo

Model details

GLM-5.3 Flash (Gonka24)

The model overview is temporarily unavailable.

LLM Gatewaygonka24/glm-5.3-flashglm-flash

Quick Info

Powered by
Provider
LLM Gateway
Model key
gonka24/glm-5.3-flash
Release date
Aug 26, 2026
Last updated
Aug 26, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.15
Output token cost
$0.30

Limits

Output tokens
16,384 tokens
Context window
200,000 tokens

Transparent token rates

Compare GLM-5.3 Flash (Gonka24) pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GLM-5.3 Flash (Gonka24)

LLM Gateway

Coverage

Z.ai published a technical account on September 17, 2026 describing how it built a production-grade inference service for GLM-5.3-Flash from scratch on a cluster of more than 100,000 Chinese-made AI accelerators, with much of the engineering work carried out by an Infra Agent powered by GLM-5.3 rather than by infrastru At the center of the account is a systems problem Z.ai calls dense feedback, which folds correctness tests, runtime logs, execution traces, runtime events, microbenchmarks, and end-to-end metrics into repeatable workflows so the agent can validate each hypothesis locally rather than waiting for a full deployment and lo

LLM Gateway

Coverage

Z.ai released GLM-5.3-Flash on August 26, 2026 as the first natively multimodal model in the GLM-5 family, shipping its weights under the MIT License. The architecture is a Mixture-of-Experts design with 320 billion total parameters but only about 18 billion active per token, a sharp reduction from the 32 billion activ On reported benchmarks, GLM-5.3-Flash scores 63.4 on DeepSWE v1.1 (up from 46.2 for GLM-5.2) and 48.8 on AutomationBench (up from 26.2), with an Artificial Analysis Intelligence Index v4.1.1 score of 57. Z.ai states the model beats GLM-5.2 across the reported coding and agentic tests at roughly one-tenth the price, and

LLM Gateway

Coverage

Z.ai confirmed that Ox Alpha, the anonymous model that had been topping OpenRouter usage charts, was in fact its own GLM-5.3-Flash, released on August 26, 2026. The model carries 320 billion total parameters with 18 billion active per token and is the first in the GLM-5 series to support multimodal input across text, i Every request to the model during its Ox Alpha preview period was served on Chinese-made AI chips, Z.ai said, with a cluster of 100,000 domestically produced accelerators powering the workload. The company did not disclose which chipmakers supplied the hardware, and Counterpoint analysts suggested the mix likely includ

Videos about GLM-5.3 Flash (Gonka24)

More models around GLM-5.3 Flash (Gonka24)