Sulat.com
AI models
OpenRouter logo

Model details

GLM Flash Latest

GLM Flash Latest is an always-redirecting entry point to the currently selected Flash-family model, which is GLM 5.3 Flash. That target is designed for efficient coding and long-running agent work, and its native multimodal design combines sparse and linear attention to support accurate long-context behavior. Applications can use the conventional OpenAI-compatible chat-completions flow, making the alias a stable integration layer while the underlying Flash model can change.

The model is a practical fit for coding assistants, document- and media-aware workflows, and agents that need to maintain context across extended task sequences. The redirect target publishes GPQA Diamond and TAU-Bench results across hosting partners, though the varying scores indicate that deployment choice can affect measured performance. Its batch endpoint is routed through a single provider and offers a separately documented cache-read rate, while the alias itself intentionally leaves pricing, routing, and benchmark details to the current target.

OpenRouter~z-ai/glm-flash-latestglm-flash

Quick Info

Powered by
Provider
OpenRouter
Model key
~z-ai/glm-flash-latest
Release date
Aug 27, 2026
Last updated
Aug 27, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.07125
Output token cost
$0.2375

Limits

Output tokens
131,072 tokens
Context window
1,310,720 tokens

Transparent token rates

Compare GLM Flash Latest pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GLM Flash Latest

OpenRouter

Official sourceBenchmark

The dedicated batch page for GLM-5.3-Flash — the current target of the Flash Latest alias — describes it as a native multimodal model from Z.ai purpose-built for efficient coding and long-horizon agent tasks, with a hybrid sparse and linear attention architecture that maintains accurate long-context behavior while redu Routing for the batch endpoint is single-provider via Together, with cache reads priced at $0.03 per 1M tokens, and the page exposes benchmarking data including GPQA Diamond and TAU-Benchvgi scores across hosting partners such as Phala (90.8% GPQA), BaseTen (84.2%), Wafer (88.3% GPQA, 80.0% TAU-Bench), DigitalOcean (89

OpenRouter

Official sourceBenchmark

The OpenRouter model page for the slug ~z-ai/glm-flash-latest confirms that GLM Flash Latest is an always-redirecting alias that currently resolves to GLM 5.3 Flash from Z.ai. The page documents the exact OpenAI-compatible integration points developers need: POST https://openrouter.ai/api/v1/chat/completions for chat, Because it is an alias, the canonical page intentionally omits per-token pricing, provider routing details, and benchmark breakdowns — those properties inherit from whatever Flash-family model is the current redirect target. The page serves as the drop-in code reference, describing itself as the entry point whose only

Videos about GLM Flash Latest

More models around GLM Flash Latest