Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Cloudflare Workers AI logo

Model details

GPT OSS 120B

GPT OSS 120B is the larger of two open-weight reasoning models released by OpenAI in the gpt-oss family, paired with a smaller 20B sibling. Both models are distributed under the Apache 2.0 license, which permits local or cloud hosting, fine-tuning, modification, and commercial use, marking OpenAI's first open-source language model release since Whisper and CLIP. The 120B variant is intended for production, general-purpose workloads that demand stronger reasoning depth than the 20B version, which is targeted at edge devices and laptops with 16 GB of memory. Through Cloudflare's hosted inference, the model becomes accessible without standing up dedicated GPU infrastructure, broadening its reach for developers building on the Workers AI platform.

Training drew on a mixture of reinforcement learning and techniques informed by OpenAI's more advanced internal systems, including o3 and other frontier models, which shaped its tool use and chain-of-thought behavior. OpenAI positions GPT OSS 120B as reaching near-parity with o4-mini on core reasoning benchmarks while running efficiently on a single 80 GB GPU, and reports strong results on tool use, few-shot function calling, and agentic evaluations such as Tau-Bench, where it reportedly outperforms proprietary peers like o1 and GPT-4o on HealthBench. These characteristics make the model a practical fit for agentic workflows, instruction-heavy applications, and teams that want open-weight flexibility combined with reasoning and function-calling competence.

Cloudflare Workers AI@cf/openai/gpt-oss-120bgpt-oss

Quick Info

Powered by
Provider
Cloudflare Workers AI
Model key
@cf/openai/gpt-oss-120b
Release date
Aug 5, 2025
Last updated
Aug 5, 2025
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.35
Output token cost
$0.75

Limits

Output tokens
16,384 tokens
Context window
128,000 tokens

Transparent token rates

Compare GPT OSS 120B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GPT OSS 120B

No articles yet. Fetch the latest news to show it here.

Videos about GPT OSS 120B

More models around GPT OSS 120B