GPT OSS Safeguard 120B is an open-weight reasoning model from OpenAI designed specifically for safety classification tasks, where it interprets a developer-provided policy at inference time to evaluate user messages, completions, and full chats. Rather than relying on a classifier trained to indirectly infer decision boundaries from large labeled datasets, the model applies chain-of-thought reasoning directly to the supplied policy, letting each developer draw their own content and behavior lines. This policy-driven approach makes it straightforward to revise rules iteratively and tune performance without retraining, which is especially useful for moderation, compliance, and content triage pipelines.
As a fine-tuned derivative of OpenAI's earlier gpt-oss open models, the 120B safeguard variant ships under the same permissive Apache 2.0 license and is downloadable from Hugging Face, making it suitable for self-hosted and customized deployments. OpenAI positions it as part of a research preview that also includes a smaller 20B sibling, giving teams a choice between deeper reasoning capacity and lighter footprint. The 120B variant is best fit for organizations that want a transparent, inspectable safety layer whose decisions can be audited through its visible chain-of-thought, rather than an opaque black-box classifier.