Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
ZenMux logo

Model details

GLM-Image

GLM-Image is an open-weights image generator from Z.ai designed to produce images that accurately depict the language rendered within them, a long-standing weak point for image synthesis models. Rather than relying on a single architecture, the system works in two coordinated stages: one stage determines the overall layout of an image while a second stage fills in the visual details. This staged design helps the model reason about where text should appear before committing to pixel-level content, which is why it has been highlighted for outperforming both open and proprietary competitors on text rendering tasks.

The architecture blends an autoregressive transformer responsible for layout planning with a diffusion-based decoder responsible for image synthesis. The autoregressive component carries around 9 billion parameters and is fine-tuned from an earlier GLM-4 language model, giving it strong grounding in textual reasoning. The diffusion decoder carries around 7 billion parameters and is built on top of an earlier CogView4 diffusion transformer, inheriting that model's visual generation capabilities. The combined system accepts text or text-plus-image prompts and produces images at resolutions ranging from 1,024 by 1,024 up to 2,048 by 2,048 pixels, making it well suited for applications such as poster design, infographics, and other graphics where legible in-image text is a core requirement.

ZenMuxz-ai/glm-image

Quick Info

Powered by
Provider
ZenMux
Model key
z-ai/glm-image
Release date
Jan 19, 2026
Last updated
Jan 19, 2026
Input modalities
Output modalities
Capabilities

Limits

Output tokens
0 tokens
Context window
10,240 tokens

Latest news about GLM-Image

ZenMux

CoverageBenchmark

LMSYS Org published a detailed technical post on August 5, 2026 describing full-stack performance optimizations for serving GLM-Image's hybrid AR+DiT pipeline in SGLang. The work replaces the Hugging Face backend with SRT to accelerate the auto-regressive stage and resolves parallelism conflicts, applying dedicated ten The team also implemented a disaggregated AR-to-DiT fan-out architecture, running one denoiser per device in parallel for the DiT stage and overlapping AR and DiT workflows by buffering AR outputs. The post frames GLM-Image as exemplifying the AR+DiT hybrid trend: a 9B vision-language model autoregressively generates s

Videos about GLM-Image