DigitalOcean
8x NVIDIA GB10 GPT OSS 120B Concurrency Vs TP. 8x NVIDIA GB10 Cluster Monitoring Under Load. 8x NVIDIA GB10 GPT OSS 120B Concurrency Vs TP.
Model details
The gpt-oss-120b model uses a mixture-of-experts architecture with 117B total parameters and 5.1B active parameters per token, enabling efficient inference on a single 80GB GPU like an NVIDIA H100 or AMD MI300X. The model was trained on OpenAI's harmony response format and is designed for agentic workflows, offering configurable reasoning effort across low, medium, and high settings. It provides full chain-of-thought visibility for debugging purposes and strong instruction-following capabilities, making it well-suited for production reasoning tasks and tool-augmented applications. The model is fine-tunable, allowing developers to customize it for specific use cases without being locked into a single deployment pattern.
This model was trained using a blend of reinforcement learning and techniques informed by OpenAI's most advanced internal systems, including o3 and other frontier models. It achieves near-parity with OpenAI o4-mini on core reasoning benchmarks while demonstrating strong performance on agentic evaluation suites like Tau-Bench and HealthBench, even outperforming proprietary models such as o1 and GPT-4o on tool use and few-shot function calling tasks. The Apache 2.0 license removes commercial restrictions, making it freely usable for experimentation, customization, and deployment. Third-party ecosystem work, such as Multiverse Computing's 50% compressed HyperNova variant, shows the model's flexibility for specialized deployment scenarios across cloud infrastructure and enterprise systems like IBM watsonx Orchestrate on Amazon Bedrock.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
DigitalOcean
8x NVIDIA GB10 GPT OSS 120B Concurrency Vs TP. 8x NVIDIA GB10 Cluster Monitoring Under Load. 8x NVIDIA GB10 GPT OSS 120B Concurrency Vs TP.
DigitalOcean
As AI agent deployments grow across enterprise systems, organizations need a way to bring together foundation models, tools, and governance frameworks to...
DigitalOcean
HyperNova 60B 2602, a 50% compressed version of OpenAI’s gpt-oss-120B, accelerates Multiverse’s plans to deliver hyper-efficient, high-performance models for free to developersDONOSTIA, Spain, Feb. 24, 2026 (GLOBE NEWSWIRE) -- Multiverse Computing, the leader in AI model compression, today announced the release of Hype
DigitalOcean
Multiverse Computing releases a compressed version of OpenAI's gpt-oss-120B
DigitalOcean
HyperNova 60B 2602, a 50% compressed version of OpenAI’s gpt-oss-120B, accelerates Multiverse’s plans to deliver hyper-efficient, high-performance models for free to developersDONOSTIA, Spain, Feb. 24, 2026 (GLOBE NEWSWIRE) -- Multiverse Computing, the leader in AI model compression, today announced the release of Hype
DigitalOcean
Amazon Bedrock Custom Model Import now supports OpenAI models with open weights, including GPT-OSS variants with 20-billion and 120-billion...
This exact model name is also listed by 3 other providers.