Synthetic
As AI agent deployments grow across enterprise systems, organizations need a way to bring together foundation models, tools, and governance frameworks to...
Model details
GPT OSS 120B is a 117B-parameter Mixture-of-Experts model that activates only 5.1B parameters per token, a design that lets it deliver high reasoning capability while running on a single 80GB GPU such as an NVIDIA H100 or AMD MI300X. OpenAI released it on August 5, 2025 alongside a smaller 21B sibling, and both models ship under the permissive Apache 2.0 license, making them freely usable for experimentation, fine-tuning, and commercial deployment. The model's training drew on reinforcement learning and techniques informed by OpenAI's more advanced internal systems, including o3, which helps explain why the 120B variant achieves reasoning performance close to o4-mini on core benchmarks despite its open-weight footprint.
Beyond raw reasoning, GPT OSS 120B is built for production agentic workflows, offering configurable reasoning effort (low, medium, high), full chain-of-thought access for debugging, native function calling, and strong tool-use behavior that has been independently measured on evaluations like Tau-Bench and HealthBench, in some cases surpassing proprietary peers. It must be used with OpenAI's harmony response format to operate correctly, and a 131k-token context window with a May 2024 knowledge cutoff makes it well suited to long-document analysis and multi-step tool orchestration. Availability has expanded well beyond OpenAI's own distribution, with hosting on platforms like Amazon Bedrock and integration into enterprise tooling such as IBM watsonx Orchestrate, positioning it as a flexible open-weight option for teams building agentic applications on their own infrastructure.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Synthetic
As AI agent deployments grow across enterprise systems, organizations need a way to bring together foundation models, tools, and governance frameworks to...
Synthetic
A peer-reviewed article published in MDPI Applied Sciences (2025) by M. Nawalny, titled "Comparative Evaluation of GPT-4o, GPT-OSS-120B and Llama-3.1-8B-Instruct Language Models in a Reproducible CV-to-JSON Extraction Pipeline," provides independent benchmark evidence on GPT-OSS-120B's performance in a structured extra The paper supplies reproducibility-focused evaluation methodology for a real-world document-understanding workload, giving technical signal about GPT-OSS-120B's relative strengths and weaknesses versus proprietary (GPT-4o) and other open-weight (Llama-3.1-8B) baselines. Its narrow scope—a single CV-to-JSON extraction t
Synthetic
Artificial Analysis publishes a third-party intelligence, performance, and price profile for OpenAI's gpt-oss-120b (high reasoning variant) that is directly useful to developers accessing the model through Synthetic's hosted endpoint. The model scores 24 on Artificial Analysis's Intelligence Index (well above the 9 ave The same profile lists concrete technical specifications that map onto Synthetic's gpt-oss-120b deployment: 131k token context window, text-in/text-out, knowledge cutoff May 31, 2024, 117B total parameters with 5.1B active per token, Apache 2.0 license, and explicit reasoning-mode support. Because Synthetic serves gpt-
This exact model name is also listed by 33 other providers.