Sulat.com
AI models
Pendra logo

Model details

GPT OSS 120B

GPT OSS 120B is the larger of two open-weight reasoning models released as part of the gpt-oss family, designed to bring frontier-style reasoning to deployments that can be freely run, modified, and fine-tuned. It was trained using a combination of reinforcement learning and techniques informed by OpenAI's more advanced internal systems, including o3. The model is built on a Mixture-of-Experts architecture with 128 experts, using a routing system that activates only the relevant experts per query, which helps balance capability with efficient inference. Its scale places it in the data center tier, where it can run on a single 80 GB GPU, and it supports a long context window suited to extended reasoning, multi-turn tasks, and document-heavy workflows.

In practice, GPT OSS 120B targets high-volume workloads and large-scale deployments where teams want strong reasoning without dependence on a closed API. The model reaches near-parity with OpenAI o4-mini on core reasoning benchmarks, while also performing strongly on tool use, few-shot function calling, and chain-of-thought evaluations such as Tau-Bench and HealthBench, where it has been reported to outperform some proprietary predecessors. Configurable reasoning effort lets users tune the balance between quality and speed, and the model is compatible with OpenAI's Responses API, making it well suited for agentic pipelines, local experimentation, and cost-sensitive production use cases that benefit from an Apache 2.0 license.

Pendragpt-oss:120bgpt-oss

Quick Info

Powered by
Provider
Pendra
Model key
gpt-oss:120b
Release date
Aug 5, 2025
Last updated
Aug 5, 2025
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
32,768 tokens
Context window
131,072 tokens

Latest news about GPT OSS 120B

Videos about GPT OSS 120B

More models around GPT OSS 120B