Nemotron 3.5 Content Safety is a small language model published by NVIDIA, packaged as a NIM container on NVIDIA NGC and built on Google's Gemma-3-4B-it base, which NVIDIA then fine-tuned on multimodal, multilingual, and reasoning-oriented content-safety datasets. The model is positioned as a compact 4B-parameter safety classifier that extends the earlier Nemotron 3 Content Safety, adding coverage for prompts, responses, and images so a single call can judge both text and visual content. Because it is released as open weights, teams can self-host it on NVIDIA-accelerated infrastructure and integrate it with common inference frameworks for real-time moderation in production pipelines.
In practice, the model is intended for developers and enterprises that need to moderate AI inputs and outputs across languages and modalities while staying inside their own governance rules. It supports a 23-category safety taxonomy, customizable policy reasoning with concise traces before each verdict, and multilingual moderation for a dozen languages out of the box, making it suitable for global deployments. An independent guardrail benchmark run by Artificial Analysis in partnership with NVIDIA placed Nemotron 3.5 Content Safety among the specialist safety classifiers evaluated for F1 score, recall, specificity, and end-to-end latency on open datasets like WildGuardTest, ToxicChat, and XSTest, giving adopters a public reference point for its balance of catching unsafe content without over-refusing safe prompts.