Currently listed through these providers:
Model details
GPT OSS 120B
gpt-oss-120b is the larger member of OpenAI's open-weight gpt-oss family, designed for production-scale reasoning and agentic workloads. It uses a Mixture-of-Experts architecture with 117B total parameters but only 5.1B active per forward pass, a sparsity pattern that lets it deliver heavyweight reasoning while keeping inference efficient. Native MXFP4 quantization lets the model run on a single H100 GPU, and additional architectural touches such as SwiGLU activations and learned attention sinks support stable long-context behavior for developer and enterprise deployments.
Practically, the model is tuned for developers who want controllable, transparent reasoning rather than a single fixed behavior profile. Users can dial reasoning depth up or down to match latency budgets, inspect the full chain of thought rather than treating it as a black box, and lean on native tool use for function calling, browsing, and structured output generation. Its text-only design and the Apache 2.0 license make it a strong fit for on-premises or private-cloud deployments where data control matters, while its position as the larger sibling to gpt-oss-20b gives teams a clear upgrade path when they need deeper reasoning without leaving the same model family.
Quick Info
Powered by- Provider
- AKI.IO
- Model key
- gpt-oss-120b
- Release date
- Aug 5, 2025
- Last updated
- Aug 5, 2025
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.15
- Output token cost
- $0.55
Limits
- Output tokens
- 32,768 tokens
- Context window
- 128,000 tokens
Transparent token rates
Compare GPT OSS 120B pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about GPT OSS 120B
Videos about GPT OSS 120B
More models around GPT OSS 120B
This exact model name is also listed by 33 other providers.