GPT OSS 120B is OpenAI's large-scale open-weight language model, released on August 5, 2025, as part of the GPT-OSS family that also includes the smaller 20B variant. It is built on a Mixture-of-Experts architecture with 117 billion total parameters and 128 experts, activating roughly 5.1 billion parameters per forward pass so that inference stays efficient despite the model's size. Native MXFP4 quantization lets it run on a single H100 GPU, making it practical for data-center deployments that need high-throughput reasoning without sprawling hardware requirements. The weights are published openly on Hugging Face under OpenAI's account, and the Apache 2.0 license allows organizations to self-host, fine-tune, modify, and commercialize the model locally or in the cloud.
The design targets high-reasoning, agentic, and general-purpose production use cases, with capabilities positioned on par with OpenAI's o4-mini tier. Configurable reasoning effort, full chain-of-thought access, and native tool use including function calling, browsing, and structured output generation make it well suited for complex multi-step workflows and agent systems. A 131K-token context window supports long-document reasoning and extended conversations, while text-in and text-out modalities keep the interface straightforward for integration into pipelines. Teams that need strong reasoning performance, the flexibility of open weights, and the ability to deploy on their own infrastructure will get the most from this model, especially in production environments where data control and customization matter.