Currently listed through these providers:
Model details
GLM 5.3 Flash Uncensored
GLM 5.3 Flash Uncensored is a community refusal-reduced variant built on Zhipu's GLM-5.3-Flash, designed to address heavy over-refusal on benign requests like copyright queries. Unlike prompt-based jailbreaks or chat-template tricks, the uncensoring is applied directly at the weight level, producing a permanent behavioral edit that loads cleanly with standard inference runtimes. The underlying architecture follows a hybrid mixture-of-experts design that combines KDA linear attention with DeepSeek-style sparse attention, allocating roughly 18 billion active parameters per token out of a 320 billion total, while retaining the multi-token-prediction head and original vision tower from the base release. This lineage gives the model both reasoning and non-reasoning operating modes without requiring custom parsers beyond the standard reasoning-format support.
In practical terms, the model is pitched toward creative writing, unrestricted roleplay, and agent workflows where refusal behavior would otherwise block legitimate use cases. Function calling, structured output, and tool use are all supported, and the model handles long contexts comfortably thanks to the sparse-attention hybrid layout. The combination of a very large total parameter pool with a comparatively modest active-per-token footprint suggests efficiency gains relative to dense counterparts at similar capability levels, making it a reasonable choice for teams that need the GLM-5.3-Flash quality profile but require fewer content-policy interruptions in their pipelines.
Quick Info
Powered by- Provider
- NanoGPT
- Model key
- z-ai/glm-5.3-flash-uncensored
- Release date
- Jul 29, 2026
- Last updated
- Aug 27, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.35
- Output token cost
- $1.40
Limits
- Input tokens
- 262,144 tokens
- Output tokens
- 32,768 tokens
- Context window
- 262,144 tokens