Sulat.com
AI models
Weights & Biases logo

Model details

gpt-oss-120b

gpt-oss-120b is an open-weight language model released by OpenAI as part of its gpt-oss family, built on a Mixture-of-Experts architecture with 117 billion total parameters and 5.1 billion activated per forward pass. Native MXFP4 quantization lets the weights run efficiently on a single H100 GPU, lowering the barrier to self-hosted or high-throughput deployments. The model is trained in OpenAI's Harmony response format and exposes configurable reasoning depth along with full chain-of-thought access, giving developers explicit control over how much deliberation happens before each answer.

Beyond text generation, gpt-oss-120b is designed for agentic and production workloads, with native support for function calling, browsing, and structured output generation that fits into pipelines requiring JSON-constrained responses or external tool orchestration. Third-party listings report strong results on graduate-level reasoning evaluations such as GPQA Diamond, alongside competitive scores on instruction-following and conversational agent benchmarks, indicating a balanced profile rather than a narrow specialist. Open weights, a long context window, and adjustable reasoning effort make it a practical fit for teams that want a transparent reasoning model they can fine-tune or self-host while still benefiting from OpenAI-aligned agentic capabilities.

Weights & Biasesopenai/gpt-oss-120bgpt-oss

Quick Info

Powered by
Provider
Weights & Biases
Model key
openai/gpt-oss-120b
Release date
Aug 5, 2025
Last updated
Aug 5, 2025
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.03
Output token cost
$0.17

Limits

Output tokens
131,072 tokens
Context window
131,072 tokens

Transparent token rates

Compare gpt-oss-120b pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about gpt-oss-120b

Weights & Biases

CoverageRelease Notes

NVIDIA NIM for Large Language Models Release 2.0.11 officially documents several deployment constraints specific to OpenAI's gpt-oss-120b. For the MXFP4 TP8 profile on NVIDIA A10G and NVIDIA A100-SXM4-40GB, the release notes require NIM_KVCACHE_PERCENT=0.80 so that the sampler warm-up has sufficient device memory. Addi The fused mixture-of-experts kernels shared by gpt-oss-120b, gpt-oss-20b, and nemotron-3-nano support LoRA adapters with a maximum rank of 128, with adapters above that rank explicitly unsupported. The NIM 2.0.11 release also updates the inference backend to vLLM 0.27.0 and ships SGLang Model-Free NIM container v2.1.2.

Videos about gpt-oss-120b

More models around gpt-oss-120b