Currently listed through these providers:
Model details
gpt-oss-safeguard-20b
The gpt-oss-safeguard-20b model is a safety-focused reasoning model built by post-training OpenAI's open-weight gpt-oss base models. Its core innovation lies in how it approaches content classification: rather than relying on a fixed set of labels baked in during training, the model interprets a developer-provided policy at inference time, allowing it to classify user messages, completions, and full conversations against whichever rules the developer defines. This architecture uses a mixture-of-experts design with 21 billion total parameters and 3.6 billion active parameters, which lets it selectively engage relevant capacity during reasoning. The model surfaces full chain-of-thought reasoning in a structured format, so developers can audit how it reaches each classification decision. Configurable reasoning effort—low, medium, or high—enables teams to trade off between speed and depth depending on the stakes of the task.
The model emerged from internal development at OpenAI before being released as a research preview under the Apache 2.0 license, with feedback from the open-source community shaping its final form. Rather than the traditional approach of training a classifier from thousands of labeled examples to indirectly infer a decision boundary, gpt-oss-safeguard is trained to reason directly from the provided policy itself, making it far more adaptable when policies evolve or need rapid iteration. Developers can revise their policies without retraining the model, and the transparency of the reasoning chain supports easier debugging and trust-building. Available through Hugging Face, Fireworks AI, LM Studio, and Groq, with fine-tuning support via LoRA on certain platforms, the model is designed for safety and content moderation pipelines rather than as a general-purpose end-user interface—OpenAI recommends the base gpt-oss models for those broader applications.
Quick Info
Powered by- Provider
- Vercel AI Gateway
- Model key
- openai/gpt-oss-safeguard-20b
- Release date
- Oct 29, 2025
- Last updated
- Dec 1, 2024
- Knowledge cutoff
- 2024-10
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.07
- Output token cost
- $0.20
Limits
- Input tokens
- 112,000 tokens
- Output tokens
- 16,000 tokens
- Context window
- 128,000 tokens
Transparent token rates
Compare gpt-oss pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about gpt-oss-safeguard-20b
No articles yet. Fetch the latest news to show it here.
Videos about gpt-oss-safeguard-20b
More models around gpt-oss-safeguard-20b
This exact model name is also listed by 3 other providers.