Sulat.com
AI models
Privatemode AI logo

Model details

gpt-oss-120b

The gpt-oss-120b is OpenAI's production-tier open-weight model designed for complex reasoning and autonomous task execution. With 117B total parameters and 5.1B active parameters during inference, the model is engineered to fit on a single 80GB GPU such as NVIDIA H100 or AMD MI300X, making high-reasoning workloads accessible without sprawling infrastructure. It is trained exclusively on OpenAI's harmony response format and requires that format to function correctly. The architecture supports configurable reasoning effort across low, medium, and high settings, allowing developers to trade off depth and latency based on their application needs, and it exposes complete chain-of-thought visibility so users can inspect and debug the model's reasoning path rather than just accepting outputs at face value.

Training blended reinforcement learning techniques with methods informed by OpenAI's frontier models including o3, aiming to transfer advanced reasoning behaviors into an open-weight package. On core reasoning benchmarks, gpt-oss-120b reaches near-parity with OpenAI o4-mini while outperforming comparably sized open models, and on agentic evaluations like Tau-Bench and HealthBench it surpassed proprietary models such as o1 and GPT-4o in tool use and function calling. Under the permissive Apache 2.0 license, organizations can fine-tune, customize, and deploy the model commercially without copyleft constraints. Its native function-calling capabilities, instruction-following strength, and Responses API compatibility make it well-suited for agentic pipelines and developer workflows where models must orchestrate tools, maintain context over long conversations, and operate reliably in production environments.

Privatemode AIgpt-oss-120bgpt-oss

Quick Info

Powered by
Provider
Privatemode AI
Model key
gpt-oss-120b
Release date
Aug 5, 2025
Last updated
Aug 5, 2025
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.4969
Output token cost
$1.9644

Limits

Output tokens
32,768 tokens
Context window
128,000 tokens

Latest news about gpt-oss-120b

Privatemode AI

CoverageRelease Notes

Discover more about what's new at AWS with OpenAI GPT OSS and NVIDIA Nemotron Models Available on Amazon Bedrock in AWS GovCloud (US)

Privatemode AI

CoverageComparison

8x NVIDIA GB10 GPT OSS 120B Throughput Vs TP Width. 8x NVIDIA GB10 Cluster Monitoring Under Load. 8x NVIDIA GB10 GPT OSS 120B Throughput Vs TP Width.

Privatemode AI

CoverageComparison

8x NVIDIA GB10 GPT OSS 120B Concurrency Vs TP. 8x NVIDIA GB10 Cluster Monitoring Under Load. 8x NVIDIA GB10 GPT OSS 120B Concurrency Vs TP.

Privatemode AI

Coverage

New prompt-based policies aim to standardise how AI systems protect under-18 users.

Privatemode AI

Coverage

As AI agent deployments grow across enterprise systems, organizations need a way to bring together foundation models, tools, and governance frameworks to...

Privatemode AI

Coverage

HyperNova 60B 2602, a 50% compressed version of OpenAI’s gpt-oss-120B, accelerates Multiverse’s plans to deliver hyper-efficient, high-performance models for free to developersDONOSTIA, Spain, Feb. 24, 2026 (GLOBE NEWSWIRE) -- Multiverse Computing, the leader in AI model compression, today announced the release of Hype

Videos about gpt-oss-120b

More models around gpt-oss-120b