Currently listed through these providers:
Model details
GPT OSS 120B
GPT OSS 120B is OpenAI's large-scale entry in its first open-weight language model family, designed to bring frontier-style reasoning into a freely deployable package. Under the hood it is a 117B-parameter Mixture-of-Experts transformer that activates only about 5.1B parameters per forward pass, with native MXFP4 quantization letting it run on a single high-end data-center GPU. The architecture is tuned for configurable chain-of-thought depth and full reasoning visibility, so developers can dial effort up or down per request while still seeing the model's thinking trace. Native function calling, web browsing, and structured output are baked in, framing the model less as a chatbot and more as an agentic engine that can chain tools together for multi-step work. The split between active and total parameters is the central design idea: keep the capacity of a very large model while paying the inference cost of a much smaller one, giving production deployments a path to reasoning quality that previously required dedicated clusters.
Released under Apache 2.0 alongside a smaller 20B sibling, GPT OSS 120B slots into the lineup as the data-center counterpart aimed at high-volume workloads, with reasoning ability compared to OpenAI's o4-mini tier. Its post-training emphasizes tool-aware behavior and controlled reasoning effort, so the same weights can behave like a quick assistant on simple prompts and a deliberate planner on hard ones, and host platforms expose tuning knobs such as minimum reasoning tokens to shape this trade-off. In practice it shows up strong on graduate-level reasoning evaluations like GPQA Diamond and on composite coding indices, while still supporting very long contexts for document- and repository-scale tasks. The combination of open weights, permissive licensing, and agentic features makes it a flexible foundation for fine-tuning, distillation into smaller deployments, and integration into retrieval or tool-use pipelines, fitting naturally into teams that want one open model across reasoning-heavy, code-heavy, and orchestration-heavy workloads.
Quick Info
Powered by- Provider
- Fireworks AI
- Model key
- accounts/fireworks/models/gpt-oss-120b
- Release date
- Aug 5, 2025
- Last updated
- Jun 16, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.15
- Output token cost
- $0.60
Limits
- Output tokens
- 32,768 tokens
- Context window
- 131,072 tokens
Transparent token rates
Compare gpt-oss pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about GPT OSS 120B
Videos about GPT OSS 120B
More models around GPT OSS 120B
This exact model name is also listed by 34 other providers.