Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Cortecs logo

Model details

GPT OSS Safeguard 120B

GPT OSS Safeguard 120B is an open-weight reasoning model from OpenAI designed specifically for safety classification tasks, where it interprets a developer-provided policy at inference time to evaluate user messages, completions, and full chats. Rather than relying on a classifier trained to indirectly infer decision boundaries from large labeled datasets, the model applies chain-of-thought reasoning directly to the supplied policy, letting each developer draw their own content and behavior lines. This policy-driven approach makes it straightforward to revise rules iteratively and tune performance without retraining, which is especially useful for moderation, compliance, and content triage pipelines.

As a fine-tuned derivative of OpenAI's earlier gpt-oss open models, the 120B safeguard variant ships under the same permissive Apache 2.0 license and is downloadable from Hugging Face, making it suitable for self-hosted and customized deployments. OpenAI positions it as part of a research preview that also includes a smaller 20B sibling, giving teams a choice between deeper reasoning capacity and lighter footprint. The 120B variant is best fit for organizations that want a transparent, inspectable safety layer whose decisions can be audited through its visible chain-of-thought, rather than an opaque black-box classifier.

Cortecsgpt-oss-safeguard-120bgpt-oss

Quick Info

Powered by
Provider
Cortecs
Model key
gpt-oss-safeguard-120b
Release date
Oct 29, 2025
Last updated
Oct 29, 2025
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.179
Output token cost
$0.697

Limits

Output tokens
128,000 tokens
Context window
128,000 tokens

Transparent token rates

Compare GPT OSS Safeguard 120B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GPT OSS Safeguard 120B

Cortecs

CoverageAnalysis

CNBC reported on October 29, 2025 that OpenAI announced two open-weight reasoning models called gpt-oss-safeguard-120b and gpt-oss-safeguard-20b, designed for developers to classify a range of online safety harms on their platforms. The models are fine-tuned, adapted versions of OpenAI's gpt-oss models released in Augu The article explains that organizations can configure the safeguard models to their specific policy needs, and because they are reasoning models that show their work, developers gain direct insight into how classifications are reached. OpenAI developed the models in partnership with Robust Open Online Safety Tools (ROO

Amazon Bedrock

Official sourceBenchmark

The full PDF technical report from OpenAI provides the detailed evaluation methodology for gpt-oss-safeguard-120b and gpt-oss-safeguard-20b, covering safety classification performance, multilingual performance, and observed challenges including disallowed content, jailbreaks, instruction hierarchy behavior, hallucinate The report documents baseline safety metrics measured in chat settings and notes that traditional classifiers trained on tens of thousands of labeled examples may still outperform these reasoning models on complex classification tasks. It explicitly directs developers to the original gpt-oss model card for underlying a

Videos about GPT OSS Safeguard 120B

More models around GPT OSS Safeguard 120B