Sulat.com
AI models
Fireworks AI logo

Model details

GPT OSS 120B

GPT OSS 120B is OpenAI's large-scale entry in its first open-weight language model family, designed to bring frontier-style reasoning into a freely deployable package. Under the hood it is a 117B-parameter Mixture-of-Experts transformer that activates only about 5.1B parameters per forward pass, with native MXFP4 quantization letting it run on a single high-end data-center GPU. The architecture is tuned for configurable chain-of-thought depth and full reasoning visibility, so developers can dial effort up or down per request while still seeing the model's thinking trace. Native function calling, web browsing, and structured output are baked in, framing the model less as a chatbot and more as an agentic engine that can chain tools together for multi-step work. The split between active and total parameters is the central design idea: keep the capacity of a very large model while paying the inference cost of a much smaller one, giving production deployments a path to reasoning quality that previously required dedicated clusters.

Released under Apache 2.0 alongside a smaller 20B sibling, GPT OSS 120B slots into the lineup as the data-center counterpart aimed at high-volume workloads, with reasoning ability compared to OpenAI's o4-mini tier. Its post-training emphasizes tool-aware behavior and controlled reasoning effort, so the same weights can behave like a quick assistant on simple prompts and a deliberate planner on hard ones, and host platforms expose tuning knobs such as minimum reasoning tokens to shape this trade-off. In practice it shows up strong on graduate-level reasoning evaluations like GPQA Diamond and on composite coding indices, while still supporting very long contexts for document- and repository-scale tasks. The combination of open weights, permissive licensing, and agentic features makes it a flexible foundation for fine-tuning, distillation into smaller deployments, and integration into retrieval or tool-use pipelines, fitting naturally into teams that want one open model across reasoning-heavy, code-heavy, and orchestration-heavy workloads.

Fireworks AIaccounts/fireworks/models/gpt-oss-120bgpt-oss

Quick Info

Powered by
Provider
Fireworks AI
Model key
accounts/fireworks/models/gpt-oss-120b
Release date
Aug 5, 2025
Last updated
Jun 16, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.15
Output token cost
$0.60

Limits

Output tokens
32,768 tokens
Context window
131,072 tokens

Transparent token rates

Compare gpt-oss pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GPT OSS 120B

Videos about GPT OSS 120B

More models around GPT OSS 120B