Currently listed through these providers:
Model details
nemotron-content-safety-reasoning-4b
Nemotron-Content-Safety-Reasoning-4B is built as a focused content safety classifier rather than a general-purpose assistant, derived from Google Gemma-3-4B-it and structured as a transformer decoder-only model with the cataloged API limit context window. Its design centers on acting as a programmable guardrail: developers can supply their own safety policies and have the model adapt its judgments and reasoning to those user-defined rules instead of relying on a fixed taxonomy. This makes it well suited for applications such as chatbots, customer support agents, and AI assistants where the definition of acceptable output shifts by deployment, brand voice, or regulatory context.
The model's standout operational feature is a dual-mode inference design that lets teams trade off latency for interpretability per request. A low-latency "Reasoning Off" mode handles straightforward or well-known policy checks efficiently, while a "Reasoning On" mode generates explicit chain-of-thought traces that improve accuracy on complex or novel policies and also provide auditable explanations for moderation decisions. Nvidia distributes the weights openly on Hugging Face and ships a deployment tutorial through the NeMo Guardrails library, while a NIM-packaged variant is offered in the Microsoft Foundry catalog, making it straightforward to integrate into both self-hosted and managed real-time moderation pipelines.
Quick Info
Powered by- Provider
- Nvidia
- Model key
- nvidia/nemotron-content-safety-reasoning-4b
- Release date
- Jan 22, 2026
- Last updated
- Jan 22, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
A provider subscription or plan supersedes token-based pricing for this model.
Limits
- Output tokens
- 4,096 tokens
- Context window
- 128,000 tokens