Model details
Qwen3Guard-Gen-8B
Built as a fine-tune of the Qwen3-8B base on roughly 1.19 million labeled safety examples, Qwen3Guard-Gen-8B takes a deliberately different approach to content moderation by treating classification as a text-generation task rather than relying on traditional classifiers. The model outputs structured safety labels and harm-type categories in natural language, which lets it apply the same workflow to both incoming user prompts and outgoing AI responses. This generative framing gives downstream applications flexible severity control while still surfacing granular harm labels such as violence, illegal acts, sexual content, personal information, self-harm, unethical behavior, political misinformation, and copyright violations. Multilingual coverage across 119 languages makes it equally useful for global products that need consistent moderation behavior across diverse user populations.
For practitioners, the model's compact 8.2-billion-parameter footprint, Apache 2.0 licensing, and non-gated Hugging Face availability make it straightforward to self-host or wire into an existing inference stack using vLLM, SGLang, or standard chat-templated transformers pipelines. Because it is open-weight, teams can audit labels, adjust category definitions, or fine-tune further on domain-specific content without depending on a proprietary moderation API. The text-generation framing also means severity decisions can be tuned through prompt design and sampling rather than fixed thresholds, which is useful when policies need to differ between prompts and responses or between different product surfaces.
Quick Info
Powered by- Provider
- OVHcloud AI Endpoints
- Model key
- qwen3guard-gen-8b
- Release date
- Jan 22, 2026
- Last updated
- Jan 22, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
A provider subscription or plan supersedes token-based pricing for this model.
Limits
- Output tokens
- 16,384 tokens
- Context window
- 32,768 tokens