Sulat.com
AI models
Get $10 off from Venice
Venice AI logo

Model details

OpenAI GPT OSS 120B

Designed as the production-oriented member of the gpt-oss lineup, this model targets high-reasoning workloads that still fit on a single 80GB accelerator such as an NVIDIA H100 or AMD MI300X, with 117B total parameters and roughly 5.1B active at inference. It is released under the permissive Apache 2.0 license, so it can be downloaded, fine-tuned, and integrated into custom pipelines without copyleft or patent friction. To work correctly, it has to be driven through the harmony response format used in training, which keeps the developer experience consistent for chat, tool calls, and structured outputs.

In practice the model behaves like a reasoning-focused generalist rather than a narrow specialist. Reasoning effort can be dialed between low, medium, and high to trade latency for depth, and the full chain-of-thought is exposed for debugging and trust without being meant for end users. Native agentic capabilities make it well suited to function calling and multi-step tool orchestration, while parameter fine-tuning lets teams specialize it for domain tasks. A long context window supports extended document analysis and multi-turn agent runs, which combined with open weights makes it a strong fit for teams that want controllable, self-hosted reasoning without sacrificing developer ergonomics.

Venice AIopenai-gpt-oss-120bgpt-oss

Quick Info

Powered by
Provider
Venice AI
Model key
openai-gpt-oss-120b
Release date
Nov 6, 2025
Last updated
Jun 11, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.07
Output token cost
$0.30

Limits

Output tokens
16,384 tokens
Context window
128,000 tokens

Latest news about OpenAI GPT OSS 120B

Venice AI

CoverageBenchmark

On August 6, 2025, OpenAI released two open-source models:

Videos about OpenAI GPT OSS 120B

More models around OpenAI GPT OSS 120B