Deep Infra
8x NVIDIA GB10 GPT OSS 120B Concurrency Vs TP. 8x NVIDIA GB10 Cluster Monitoring Under Load. 8x NVIDIA GB10 GPT OSS 120B Concurrency Vs TP.
Model details
GPT OSS 120B is OpenAI's large-scale open-weight language model, released on August 5, 2025, as part of the GPT-OSS family that also includes the smaller 20B variant. It is built on a Mixture-of-Experts architecture with 117 billion total parameters and 128 experts, activating roughly 5.1 billion parameters per forward pass so that inference stays efficient despite the model's size. Native MXFP4 quantization lets it run on a single H100 GPU, making it practical for data-center deployments that need high-throughput reasoning without sprawling hardware requirements. The weights are published openly on Hugging Face under OpenAI's account, and the Apache 2.0 license allows organizations to self-host, fine-tune, modify, and commercialize the model locally or in the cloud.
The design targets high-reasoning, agentic, and general-purpose production use cases, with capabilities positioned on par with OpenAI's o4-mini tier. Configurable reasoning effort, full chain-of-thought access, and native tool use including function calling, browsing, and structured output generation make it well suited for complex multi-step workflows and agent systems. A 131K-token context window supports long-document reasoning and extended conversations, while text-in and text-out modalities keep the interface straightforward for integration into pipelines. Teams that need strong reasoning performance, the flexibility of open weights, and the ability to deploy on their own infrastructure will get the most from this model, especially in production environments where data control and customization matter.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Deep Infra
8x NVIDIA GB10 GPT OSS 120B Concurrency Vs TP. 8x NVIDIA GB10 Cluster Monitoring Under Load. 8x NVIDIA GB10 GPT OSS 120B Concurrency Vs TP.
QVAC
An arXiv preprint dated December 9, 2025 benchmarks OpenAI's GPT-OSS-120B and GPT-OSS-20B across ten financial NLP tasks, reporting 66.5% accuracy for the 120B variant against datasets such as Financial PhraseBank, FiQA-SA, and FLARE FINER-ORD. The evaluation introduces a Token Efficiency Score and tokens-per-second th The study finds a counter-intuitive efficiency paradox in which the smaller GPT-OSS-20B matches the 120B's accuracy at roughly 65.1% while delivering higher throughput and a 198.4 Token Efficiency Score, suggesting architectural and training innovations in GPT-OSS allow compact variants to compete with larger models li
Vercel AI Gateway
Epoch AI's model registry entry for gpt-oss-120b confirms OpenAI as the creator and August 5, 2025 as the release date for this open-weights model. The page reports 1.2e11 parameters, training on NVIDIA H100 SXM5 80GB hardware, and roughly 4.9e24 FLOP of training compute. It assigns an Epoch Capability Index (ECI) scor The Epoch scorecard provides concrete benchmark deltas versus GPT-4.1, o3-mini, and Qwen3-Max across agent, conceptual reasoning, games, mathematics, science, and software engineering domains — for example, 96% on GPQA Diamond, 88% on Aider Polyglot, 93% on WeirdML v2, and 100% on OTIS Mock AIME 2024-2025. These figure
Neon
BenchLM.ai, a third-party benchmark aggregator, provides a capability snapshot for GPT-OSS 120B with data current as of September 4, 2026. The model was released August 5, 2025, features open weights and a 128K context window, and scores 49.3 out of 100, ranking 137 of 232 tracked models. Its strongest eligible categor The page also reports a speed of approximately 150 tokens per second with a 14.18-second time to first token, and notes that published weights can be self-hosted, though infrastructure cost varies and no comparable first-party API token rate is published. This aggregator-derived data provides a current independent tech
This exact model name is also listed by 28 other providers.