Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Helicone logo

Model details

OpenAI GPT-OSS 120b

GPT-OSS 120b is part of OpenAI's gpt-oss family of open-weight language models, released alongside the smaller gpt-oss-20b variant under an Apache-2.0 license with an additional usage policy. The release is positioned as a step from closed to hybrid open systems, giving developers direct access to a reasoning-oriented model that can be self-hosted on suitable hardware or served through third-party inference platforms. Deep integration with ecosystems such as Hugging Face, vLLM, Ollama, LM Studio, Amazon Bedrock, and Databricks makes the model flexible for a range of deployment paths, from local experimentation to managed cloud hosting, including availability as a model-as-a-service offering on platforms like Google's Gemini Enterprise Agent Platform.

The model is purpose-built for reasoning workloads, with reported benchmark results showing it reaching 90% on MMLU and outperforming OpenAI's o4-mini on AIME math competitions and the HealthBench health-conversation benchmark. Through MXFP4 quantization, the 120b-parameter configuration is engineered to fit on roughly 80–96 GB of VRAM, bringing frontier-style reasoning to a hardware footprint that is approachable for serious on-premise setups. It also introduces the Harmony response format, a multi-channel message structure that carries chain-of-thought traces and tool calls, which aids debugging and trust in the model's intermediate reasoning. These characteristics make GPT-OSS 120b a practical fit for teams that need transparent, locally controllable reasoning for math, health, and general-purpose language applications.

Heliconegpt-oss-120bgpt-oss

Quick Info

Powered by
Provider
Helicone
Model key
gpt-oss-120b
Release date
Jun 1, 2024
Last updated
Jun 1, 2024
Knowledge cutoff
2024-06
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.04
Output token cost
$0.16

Limits

Output tokens
131,072 tokens
Context window
131,072 tokens

Transparent token rates

Compare OpenAI GPT-OSS 120b pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about OpenAI GPT-OSS 120b

Videos about OpenAI GPT-OSS 120b

More models around OpenAI GPT-OSS 120b