GPT OSS Safeguard 20B is an open-weight reasoning model post-trained from the underlying GPT-OSS family and released by OpenAI under the Apache 2.0 license alongside the GPT-OSS usage policy. It is text-only and was designed specifically to reason from a provided policy in order to label content against that policy, rather than to serve as a general-purpose chat assistant. The technical report positions the larger sibling and this 20B variant as customizable classifiers that expose full chain-of-thought reasoning, support configurable reasoning effort levels (low, medium, and high), and integrate with Structured Outputs for downstream pipelines.
On the policy-classification benchmarks reported by OpenAI, the model shows the qualitative strengths its creators highlight: it outperforms GPT-5-thinking on multi-policy accuracy and slightly edges out OpenAI's internal Safety Reasoner on the 2022 OpenAI Moderation evaluation, while remaining a compact 20B option that the report describes as preferable for moderation tasks. Multilingual performance on MMMLU is reported to track closely with the base GPT-OSS models across all reasoning effort levels, indicating that the safety-focused post-training did not regress language coverage. Practical fit centers on internal moderation and content-labeling systems where a deployable open-weight reasoner that can follow a supplied rubric is more valuable than a broad conversational model.