Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Databricks logo

Model details

GLM-5.2

GLM-5.2 is positioned as a flagship model purpose-built for long-horizon work, representing a meaningful step forward from its predecessor GLM-5.1. Its central design goal is making long-context engineering genuinely usable: sustaining quality across sprawling, messy coding-agent trajectories rather than merely accepting more input. To support that, the model introduces a new architectural pattern called IndexShare, which reuses the same indexer across every four sparse-attention layers, trimming per-token compute by a factor of 2.9 at the cataloged API limit lengths. The accompanying improvements to the multi-token prediction layer lift speculative-decoding acceptance length by up to 20%, which helps balance latency against thoroughness on extended agent runs.

For practical deployment, GLM-5.2 leans into flexible effort control, letting developers tune thinking depth to match the responsiveness they need for coding assistance or longer agentic workflows. It ships under an MIT open-source license with no regional restrictions, and the weights are openly published on HuggingFace alongside the accompanying code repository, making it straightforward to self-host or fine-tune. The combination of a robust the cataloged API limit context, sparse attention efficiency, and open availability points to strong fit for teams running sustained multi-step code generation, retrieval-heavy analysis, or research workflows where maintaining coherence over very long sessions is the primary requirement.

Databricksdatabricks-glm-5-2glm

Quick Info

Powered by
Provider
Databricks
Model key
databricks-glm-5-2
Release date
Jun 13, 2026
Last updated
Jun 13, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$1.40
Output token cost
$4.40

Limits

Output tokens
131,072 tokens
Context window
1,000,000 tokens

Transparent token rates

Compare GLM-5.2 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GLM-5.2

Databricks

Coverage

The NIST Center for AI Standards and Innovation (CAISI) published a public assessment of Z.ai's GLM-5.2 on July 8, 2026 (posted July 17, 2026), evaluating the model that Z.ai released as an open-weight release on June 16, 2026. CAISI concluded that GLM-5.2 was probably the most capable open-weight AI model at its relea CAISI also reported mixed safeguards: GLM-5.2's safeguards allow assistance with agentic cyber exploit development and block fewer sensitive biological questions than reference U.S. models, but the model appears potentially more robust against agent hijacking and jailbreaking attacks than other evaluated PRC open-weigh

Databricks

CoverageRelease Notes

Featherless served as a Day Zero launch partner when Z.ai released GLM-5.2 on June 16, 2026, providing the model on its platform from day one via an MIT-licensed OpenAI-compatible API. The model is roughly 753B parameters in a Mixture-of-Experts design activating around 39B parameters per token, and it is the highest-r Compared to GLM-5.1, GLM-5.2 shows substantial gains across coding and reasoning benchmarks: Terminal-Bench 2.1 rose from 63.5 to 81.0, SWE-bench Pro from 58.4 to 62.1, FrontierSWE from 30.5 to 74.4, and SWE-Marathon from 1.0 to 13.0, while AIME 2026 improved from 95.3 to 99.2 and GPQA-Diamond from 86.2 to 91.2. Archit

Databricks

CoverageBenchmark

On June 16, 2026, Z.ai (formerly Zhipu AI) released GLM-5.2 as a 753-billion-parameter open-weights large language model with a 1-million-token context window, targeting long-horizon autonomous coding and engineering tasks. The model's weights were released under an MIT license and made available on Hugging Face, the Z GLM-5.2 introduces an architectural optimization called IndexShare, which reuses a single indexer across every four sparse-attention layers, reducing per-token compute FLOPs by 2.9× at the maximum 1-million-token context length. The model also features an upgraded Multi-Token Prediction (MTP) layer that boosts speculat

Databricks

CoverageBenchmark

Semgrep published a security research post on June 22, 2026, reporting that GLM 5.2 outperformed Claude Opus 4.8 on Semgrep's IDOR benchmark when models were given nothing but a prompt. Among open-weight options tested against the same dataset used to evaluate frontier coding agents, GLM 5.2 produced the best results, The write-up positions GLM 5.2 as a strong open-weight contender for cyber-focused tasks, beating a frontier Claude model in a head-to-head prompt-only comparison. This is a third-party security-vendor benchmark rather than an official release note from Z.ai or Databricks, but it provides concrete comparative capabilit

Videos about GLM-5.2

More models around GLM-5.2