Currently listed through these providers:
Model details
GPT OSS 120B
GPT-OSS 120B is an open-weight language model built around a Mixture-of-Experts architecture that activates a targeted portion of its total parameters during each forward pass, keeping inference lightweight enough to run on a single 80 GB GPU while still leveraging the full 117B parameter model. Its design centers on agentic workflows: native function calling, browsing, and structured output generation let it plug directly into enterprise toolchains, while configurable reasoning effort lets developers tune how deeply the model thinks through a problem based on latency or quality needs. Full chain-of-thought access means the model's internal reasoning is exposed for debugging and audit—useful in regulated industries where trust in outputs matters. The permissive Apache 2.0 license removes barriers to commercial deployment, fine-tuning, and local hosting, making this a model both large enterprises and independent developers can build on without licensing overhead.
The model's training blended reinforcement learning with techniques informed by OpenAI's frontier systems, including o3, producing a model that achieves near-parity with proprietary reasoning models on core benchmarks while running at a fraction of the compute cost. On agentic evaluation suites like Tau-Bench and HealthBench, it demonstrated strong tool use and few-shot function calling, in some cases outperforming earlier proprietary models. Adoption is already spreading across enterprise stacks—IBM watsonx Orchestrate supports it on Amazon Bedrock, and cluster configurations like 8x NVIDIA GB10 deployments are enabling higher-concurrency production workloads. For teams building AI agents that need reliable tool use, flexible deployment options, and transparent reasoning without the vendor lock-in of closed models, GPT-OSS 120B positions itself as a practical open-weight choice.
Quick Info
Powered by- Provider
- Vertex
- Model key
- openai/gpt-oss-120b-maas
- Release date
- Aug 5, 2025
- Last updated
- Aug 5, 2025
- AI SDK package
@ai-sdk/openai-compatible- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.09
- Output token cost
- $0.36
Limits
- Output tokens
- 131,072 tokens
- Context window
- 131,072 tokens
Transparent token rates
Compare GPT OSS 120B pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about GPT OSS 120B
No articles yet. Fetch the latest news to show it here.
Videos about GPT OSS 120B
More models around GPT OSS 120B
This exact model name is also listed by 28 other providers.