Sulat.com
AI models
evroc logo

Model details

GPT OSS 120B

The gpt-oss-120b is a Mixture-of-Experts model designed to deliver frontier-level reasoning while remaining practical for resource-constrained environments. With 117 billion total parameters and 5.1 billion active parameters per forward pass, the architecture incorporates SwiGLU activations and learned attention sinks to improve inference efficiency. The model supports configurable reasoning effort, allowing users to dial in the depth of thought based on latency needs. Full chain-of-thought transparency is built in, making reasoning steps inspectable rather than opaque. Released under the permissive Apache 2.0 license, it is fully fine-tunable and optimized to run on a single 80GB GPU, making it viable for on-premises or private cloud deployments where data sovereignty is a priority.

Training blended reinforcement learning with techniques informed by OpenAI's most advanced internal systems, including o3 and other frontier models. The approach enabled the model to achieve near-parity with OpenAI o4-mini on core reasoning benchmarks while excelling at agentic tasks. On the Tau-Bench and HealthBench evaluations, gpt-oss-120b outperformed proprietary systems like OpenAI o1 and GPT-4o in tool use and function calling, demonstrating strong few-shot and chain-of-thought capabilities. Built on the harmony response format and compatible with the Responses API, the model is positioned for production agentic workflows, enterprise reasoning tasks, and developers who need an openly deployable foundation model without vendor lock-in.

evrocopenai/gpt-oss-120bgpt-oss

Quick Info

Powered by
Provider
evroc
Model key
openai/gpt-oss-120b
Release date
Aug 5, 2025
Last updated
Aug 5, 2025
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.23
Output token cost
$0.92

Limits

Output tokens
65,536 tokens
Context window
65,536 tokens

Latest news about GPT OSS 120B

evroc

CoverageComparison

8x NVIDIA GB10 GPT OSS 120B Concurrency Vs TP. 8x NVIDIA GB10 Cluster Monitoring Under Load. 8x NVIDIA GB10 GPT OSS 120B Concurrency Vs TP.

evroc

Coverage

As AI agent deployments grow across enterprise systems, organizations need a way to bring together foundation models, tools, and governance frameworks to...

evroc

Coverage

by Niithiyn Vijeaswaran, June Won, Pradyun Ramadorai, Saurabh Trikande, Breanne Warner, and Yotam Moss on 05 AUG 2025 in Amazon SageMaker, Amazon SageMaker...

evroc

Coverage

HyperNova 60B 2602, a 50% compressed version of OpenAI’s gpt-oss-120B, accelerates Multiverse’s plans to deliver hyper-efficient, high-performance models for free to developersDONOSTIA, Spain, Feb. 24, 2026 (GLOBE NEWSWIRE) -- Multiverse Computing, the leader in AI model compression, today announced the release of Hype

Videos about GPT OSS 120B

More models around GPT OSS 120B