Sulat.com
AI models
Synthetic logo

Model details

GPT OSS 120B

GPT OSS 120B is a 117B-parameter Mixture-of-Experts model that activates only 5.1B parameters per token, a design that lets it deliver high reasoning capability while running on a single 80GB GPU such as an NVIDIA H100 or AMD MI300X. OpenAI released it on August 5, 2025 alongside a smaller 21B sibling, and both models ship under the permissive Apache 2.0 license, making them freely usable for experimentation, fine-tuning, and commercial deployment. The model's training drew on reinforcement learning and techniques informed by OpenAI's more advanced internal systems, including o3, which helps explain why the 120B variant achieves reasoning performance close to o4-mini on core benchmarks despite its open-weight footprint.

Beyond raw reasoning, GPT OSS 120B is built for production agentic workflows, offering configurable reasoning effort (low, medium, high), full chain-of-thought access for debugging, native function calling, and strong tool-use behavior that has been independently measured on evaluations like Tau-Bench and HealthBench, in some cases surpassing proprietary peers. It must be used with OpenAI's harmony response format to operate correctly, and a 131k-token context window with a May 2024 knowledge cutoff makes it well suited to long-document analysis and multi-step tool orchestration. Availability has expanded well beyond OpenAI's own distribution, with hosting on platforms like Amazon Bedrock and integration into enterprise tooling such as IBM watsonx Orchestrate, positioning it as a flexible open-weight option for teams building agentic applications on their own infrastructure.

Synthetichf:openai/gpt-oss-120bgpt-oss

Quick Info

Powered by
Provider
Synthetic
Model key
hf:openai/gpt-oss-120b
Release date
Aug 5, 2025
Last updated
Aug 5, 2025
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.10
Output token cost
$0.10

Limits

Output tokens
32,768 tokens
Context window
131,072 tokens

Transparent token rates

Compare gpt-oss pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GPT OSS 120B

Synthetic

Coverage

As AI agent deployments grow across enterprise systems, organizations need a way to bring together foundation models, tools, and governance frameworks to...

Synthetic

Coverage

A peer-reviewed article published in MDPI Applied Sciences (2025) by M. Nawalny, titled "Comparative Evaluation of GPT-4o, GPT-OSS-120B and Llama-3.1-8B-Instruct Language Models in a Reproducible CV-to-JSON Extraction Pipeline," provides independent benchmark evidence on GPT-OSS-120B's performance in a structured extra The paper supplies reproducibility-focused evaluation methodology for a real-world document-understanding workload, giving technical signal about GPT-OSS-120B's relative strengths and weaknesses versus proprietary (GPT-4o) and other open-weight (Llama-3.1-8B) baselines. Its narrow scope—a single CV-to-JSON extraction t

Synthetic

CoverageBenchmark

Artificial Analysis publishes a third-party intelligence, performance, and price profile for OpenAI's gpt-oss-120b (high reasoning variant) that is directly useful to developers accessing the model through Synthetic's hosted endpoint. The model scores 24 on Artificial Analysis's Intelligence Index (well above the 9 ave The same profile lists concrete technical specifications that map onto Synthetic's gpt-oss-120b deployment: 131k token context window, text-in/text-out, knowledge cutoff May 31, 2024, 117B total parameters with 5.1B active per token, Apache 2.0 license, and explicit reasoning-mode support. Because Synthetic serves gpt-

Videos about GPT OSS 120B

More models around GPT OSS 120B