Sulat.com
AI models
Fireworks AI logo

Model details

GLM 5.3 Flash

GLM-5.3-Flash is positioned as the budget-friendly sibling within Z.ai's lineup, built around a mixture-of-experts design that pairs roughly 321B aggregate parameters with around 18B active per token. That sparsity is the central engineering idea: most of the heavy lifting is offloaded to rarely used experts, so each forward pass behaves like a much smaller model. Z.ai released the system under the preview codename "Ox Alpha" before publicly naming it, which gave the cost-optimized tier a deliberate identity separate from the flagship. The architecture is also natively multimodal, accepting text, images, video, and PDF inputs from the start rather than bolting vision on later, which matters for document-heavy and screen-aware workflows.

The practical story is competitive coding and agent behavior at a low operating cost. On Terminal-Bench 2.1 the model reaches 84.3, sitting close to leading frontier systems like Claude Opus 4.8 at 85.0 and GPT-5.6 Terra at 87.4, which is a striking result for a budget-tier release. Tool calling and structured output support make it well suited to agent loops, code generation, and retrieval-heavy assistants where low per-token cost compounds over long sessions. The very large context window lets teams point the model at whole repositories or long-running transcripts without aggressive chunking. Overall it fits teams that want near-frontier agentic quality but need to keep spend under control, including local serving through Ollama's cloud tag for prototyping and evaluation.

Fireworks AIaccounts/fireworks/models/glm-5p3-flashglm

Quick Info

Powered by
Provider
Fireworks AI
Model key
accounts/fireworks/models/glm-5p3-flash
Release date
Aug 26, 2026
Last updated
Aug 26, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.15
Output token cost
$0.50

Limits

Output tokens
131,072 tokens
Context window
1,000,000 tokens

Latest news about GLM 5.3 Flash

Videos about GLM 5.3 Flash

Recent tweets and retweets from Fireworks AI

More models around GLM 5.3 Flash