Sulat.com
AI models
submodel logo

Model details

GPT OSS 120B

GPT OSS 120B is built as a Mixture-of-Experts architecture, distributing 117 billion total parameters across specialized experts while activating only 5.1 billion per forward pass. This design choice keeps the model efficient enough to run on a single 80 GB GPU while preserving the depth needed for serious reasoning workloads. The architecture incorporates SwiGLU activations and learned attention sinks, supporting chain-of-thought processing and configurable reasoning effort levels that developers can dial up or down depending on latency needs. It was engineered specifically for agentic workflows, combining strong instruction following with native tool use capabilities including function calling, browsing, and structured output generation—all under the permissive Apache 2.0 license that removes commercial barriers for enterprise and on-premises deployment.

The model's training combined reinforcement learning with techniques borrowed from OpenAI's most advanced internal systems, including o3 and other frontier models. This lineage shows up in its benchmark performance: GPT OSS 120B achieves near-parity with o4-mini on core reasoning tasks and outperforms proprietary models like o1 and GPT-4o on agentic evaluations such as Tau-Bench and HealthBench. Third-party developers have already compressed it further for specialized applications, demonstrating the flexibility of the base weights. The full chain-of-thought visibility lets developers trace reasoning paths for debugging, while its fine-tunability opens doors for customization across industries—from healthcare compliance to customer automation—without sacrificing the open-weight transparency that makes it suitable for regulated environments.

submodelopenai/gpt-oss-120bgpt-oss

Quick Info

Powered by
Provider
submodel
Model key
openai/gpt-oss-120b
Release date
Aug 23, 2025
Last updated
Aug 23, 2025
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.10
Output token cost
$0.50

Limits

Output tokens
32,768 tokens
Context window
131,072 tokens

Latest news about GPT OSS 120B

submodel

Coverage

As AI agent deployments grow across enterprise systems, organizations need a way to bring together foundation models, tools, and governance frameworks to...

Videos about GPT OSS 120B

More models around GPT OSS 120B