Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Requesty logo

Model details

glm-5.2-fast

GLM-5.2-fast is positioned as a long-context workhorse in Zhipu AI's GLM family, purpose-built to plan, execute, and iterate over extended, engineering-grade tasks. It brings the cataloged API limit context window paired with multi-effort coding capabilities that target long-horizon development scenarios where many files, references, and prior decisions must remain in view at once. The model's design leans into agent-style behavior rather than single-prompt answering, making it a natural fit for autonomous coding loops, repository-scale refactors, and complex debugging chains that unfold over many turns.

At 743B parameters, GLM-5.2-fast pairs raw scale with two efficiency-oriented architectural moves: a new IndexShare structure and an improved multi-token prediction layer. Together they reduce per-token FLOPs and lengthen the speculative decoding window, so the model can sustain long outputs without paying a full forward pass for every token. The combination of reasoning and tool calling further reinforces its agent profile, allowing it to call external systems while maintaining coherent planning over very large contexts, which is well suited to teams running sustained, multi-step engineering workloads on cost-efficient infrastructure.

Requestyglm-5.2-fastglm

Quick Info

Powered by
Provider
Requesty
Model key
glm-5.2-fast
Release date
Jul 13, 2026
Last updated
Jul 13, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$2.10
Output token cost
$6.60

Limits

Output tokens
131,072 tokens
Context window
1,000,000 tokens

Transparent token rates

Compare glm-5.2-fast pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about glm-5.2-fast

Requesty

Official sourceBenchmark

Requesty lists a managed Fireworks AI deployment of Z.ai's GLM-5.2 as `fireworks/glm-5.2-fast`, exposed through its OpenAI-compatible router at `https://router.requesty.ai/v1`. The catalog page documents a 1M-token context window with a 131,072-token max output, chat API type, US serving region, no data retention, no t Provider rates for this Fireworks-hosted GLM-5.2 deployment on Requesty are $2.10 per 1M input tokens and $6.60 per 1M output tokens, with a cached-input rate of $0.21 per 1M and a 3.1x output-to-input ratio. Sample workload figures on the page put 100K input + 10K output at $0.28 and 10M input + 1M output at $27.60 be

Videos about glm-5.2-fast

More models around glm-5.2-fast