Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
SiliconFlow logo

Model details

GLM-5.2

GLM-5.2 is Z.ai's flagship model purpose-built for long-horizon tasks, representing a substantial leap over its predecessor GLM-5.1. It ships as a fully open-weights release under an MIT license with no regional access limits, initially appearing to coding-plan subscribers before the full weights became publicly downloadable. The model is structured as a 753-billion-parameter Mixture-of-Experts architecture with around 40 billion active parameters, and it sustains a solid 1-million-token context window designed for stable operation across extended agent trajectories rather than merely accepting very long inputs.

Z.ai highlights several technical improvements in GLM-5.2, including an architectural feature called IndexShare that reuses the same indexer across every four sparse attention layers to cut per-token compute at long context lengths, alongside refinements to the model's multi-token prediction layer that boost speculative decoding acceptance. The release emphasizes advanced coding capabilities with configurable thinking effort levels so developers can trade latency for performance, and independent coverage notes the model ranks competitively on agentic front-end coding leaderboards. GLM-5.2 is particularly well suited to project-level software engineering, long-running coding agents that must retain engineering context through multi-step workflows, and complex automation pipelines that demand consistent tool use over extended sessions.

SiliconFlowzai-org/GLM-5.2glm

Quick Info

Powered by
Provider
SiliconFlow
Model key
zai-org/GLM-5.2
Release date
Jun 13, 2026
Last updated
Jun 13, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$1.302
Output token cost
$4.092

Limits

Output tokens
262,000 tokens
Context window
1,049,000 tokens

Transparent token rates

Compare GLM-5.2 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GLM-5.2

SiliconFlow

Official sourceAnnouncement

SiliconFlow published a developer-focused comparison titled "Best Cheap API for SillyTavern: DeepSeek-V4-Flash, GLM-5.2, and Long-Context Cost" on August 13, 2026. The article frames GLM-5.2 as one of two paths for long SillyTavern sessions: a low-cost option (DeepSeek-V4-Flash) versus a model designed for more demandi The piece also explains how context growth affects token usage in SillyTavern, listing elements such as character card, scenario, example dialogue, chat history, World Info/lorebook, user persona, author's note, and model output. It notes that a 1049K context model charges only for tokens actually processed and generat

SiliconFlow

Official sourceAnnouncement

SiliconFlow's launch post "GLM-5.2 Now on SiliconFlow: 1M Context, Long-Horizon Engineering, Near-Opus Coding" (June 24, 2026) announces GLM-5.2 as Z.ai's latest open-source flagship going live on SiliconFlow. The post cites frontier-level coding and agentic performance, claiming GLM-5.2 matches or beats GPT-5.5 on ben At launch, SiliconFlow listed GLM-5.2 with a 1,049K-token context window, $1.40/M input, $4.40/M output, and $0.26/M cache read pricing, served via FP8 precision with function calling, context caching, thinking, and dual reasoning effort (High and Max) capabilities. The post highlights OpenAI- and Anthropic-compatible

SiliconFlow

CoverageBenchmark

A third-party guide updated September 9, 2026, separates independently verified GLM-5.2 benchmark results from Zhipu's self-reported numbers across categories. GLM-5.2's clearest independently verified result is its #1 finish on Design Arena's Code Categories leaderboard, a blind human-preference test where it ranks ro The guide notes that specific point scores on SWE-bench Pro (62.1), Terminal-Bench 2.1, and FrontierSWE come from Zhipu's own technical report and model card — directionally corroborated by independent evaluators but not reproduced exactly by a third party using Zhipu's precise methodology. The headline framing of 'nea

SiliconFlow

CoverageBenchmark

OpenRouter's GLM 5.2 directory page describes the model as a large-scale reasoning model from Z.ai supporting text input/output with a 1M-token context window, suited to long-horizon agent workflows, project-level software engineering, and complex multi-step automation. It notes that reasoning efforts "high" and "xhigh The page catalogs 20+ hosting providers with per-million-token input, output, and cache-read pricing, plus latency, throughput, and uptime metrics, and lets users filter by quantization and routing mode (Balanced, Nitro, Exacto). SiliconFlow appears with a 15% discount at $1.19 input / $3.74 output / $0.221 cache read

SiliconFlow

CoverageBenchmark

pricepertoken.com's GLM 5.2 API Pricing 2026 page (last updated September 1, 2026) reports that GLM 5.2 was released on June 16, 2026, with pricing starting at $0.402 per million input tokens and $1.26 per million output tokens, plus a $0.060 per 1M cached token rate. The model is listed with a 1,048,576-token context The page highlights a 71.3% drop in GLM 5.2's cheapest input price over the past 90 days, falling from $1.40 to $0.402 per million tokens, signaling aggressive price competition across providers. It lists 20 providers for the model, with FlexAI cited as the cheapest at $0.402 per million input tokens and input prices r

Together AI

CoverageBenchmark

Z.ai (formerly Zhipu AI) released GLM-5.2 on June 16, 2026, as a 753-billion-parameter open-weights large language model engineered for long-horizon autonomous coding and engineering tasks, according to VentureBeat reporting. The model ships under an MIT open-source license on Hugging Face, is also available via the Z. Architecturally, GLM-5.2 introduces an optimization called IndexShare, which reuses a single indexer across every four sparse attention layers, reducing per-token compute FLOPs by approximately 2.9x at the maximum 1-million-token context length. It also features an upgraded Multi-Token Prediction layer for speculative

Together AI

Coverage

NIST's Center for AI Standards and Innovation published an assessment of Z.ai's GLM-5.2 on July 8, 2026, reporting it was probably the most capable open-weight model at its June 16, 2026 release. CAISI found GLM-5.2's overall capabilities are similar to GPT-5.2 (December 2025) and its cyber capabilities are similar to Anthropic's Opus 4.6 (February 2026). The evaluation also flagged mixed safeguards performance: GLM-5.2 allowed assistance with agentic cyber exploit development, blocked fewer sensitive biological questions than reference U.S. models, but appeared potentially more robust against agent hijacking and jailbreaking than other evaluated PRC open-weight models. The report notes that prompt-based safeguards can still be circumvented when the open-weight model is self-hosted, a key caveat for developers.

SiliconFlow

CoverageBenchmark

An aggregator page presents all 17 published GLM-5.2 benchmark results grouped into reasoning, coding, and agentic categories, explicitly disclaiming that the figures are republished as published by the model authors and that the site does not run the evaluations. The long-horizon trio — FrontierSWE (74.4%, up to 20-ho Generation-over-generation deltas versus GLM-5.1 show uneven gains: +3.7 on SWE-bench Pro (62.1 vs 58.4), +17.5 on Terminal-Bench 2.1 (81.0 vs 63.5), +28.2 (2.6×) on DeepSWE (46.2 vs 18.0), +12.8 on ProgramBench (63.7 vs 50.9), +5.0 on MCP-Atlas, +7.5 on Tool-Decathlon (48.2 vs 40.7), and +9.5 on HLE (40.5 vs 31.0). Th

SiliconFlow

CoverageBenchmark

Artificial Analysis's "GLM-5.2 (Non-reasoning): API Provider Performance Benchmarking & Price Analysis" evaluates the model across 8 API providers. Output speed rankings put Baseten (FAST) at 215.9 t/s first, Modular at 165.5 t/s second, Nebius (FP4) at 118.9 t/s third, Baseten at 105.3 t/s fourth, and SiliconFlow (FP8 For pricing, Bitdeer AI leads at $0.37 blended per 1M tokens, followed by Wafer at $0.79, Baseten at $0.82, SiliconFlow (FP8) at $0.85, and Modular at $0.90, with up to a 6.8x spread across providers. The analysis also breaks down separate cache hit, input, and output prices per provider and notes that cache write and

Videos about GLM-5.2

More models around GLM-5.2