Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Tempr Gateway logo

Model details

GPT OSS 120B

The model overview is being prepared.

Tempr Gatewaygroq/openai/gpt-oss-120bgpt-oss

Quick Info

Powered by
Provider
Tempr Gateway
Model key
groq/openai/gpt-oss-120b
Release date
Aug 5, 2025
Last updated
Oct 21, 2025
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.15
Output token cost
$0.60

Limits

Output tokens
65,536 tokens
Context window
131,072 tokens

Transparent token rates

Compare GPT OSS 120B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GPT OSS 120B

Tempr

Coverage

OpenAI released two open-weight language models, gpt-oss-120b and gpt-oss-20b, on August 5, 2025 — its first open-weight releases since GPT-2 in 2019. The text-only models are positioned as lower-cost, customizable options aimed at developers, researchers, and enterprises, positioning OpenAI against Meta, Microsoft-backed Mistral, and DeepSeek in the open-weight space. OpenAI President Greg Brockman framed the launch as a contribution to an emerging open-weight ecosystem rather than a standalone source release. OpenAI collaborated with Nvidia, AMD, Cerebras, and Groq to validate that the gpt-oss models run efficiently across heterogeneous chip architectures, with Nvidia CEO Jensen Huang calling the move a step forward for open-source innovation. The company emphasized extensive safety training and testing on the open-weight releases before launch. The models are open-weight rather than fully open-source, meaning parameters are public but the training source code is not.

Tempr

CoverageBenchmark

Independent evaluators at 16x Engineer tested gpt-oss-120b on August 7, 2025 across four inference providers accessed via OpenRouter: Cerebras, Fireworks, Together, and Groq. The evaluation used matched writing and coding tasks under default provider settings to isolate host-side behavior. Results showed Cerebras and Groq delivering substantially faster responses than Fireworks and Together, with response-length variance attributed to the model itself rather than providers. Quality ratings on writing outputs clustered around 7.5–8 out of 10 across providers, with Cerebras, Fireworks, and Together producing at least one strong response and Groq consistently scoring 7.5. The study corroborates OpenAI's Romain Huet's acknowledgment that performance and correctness vary across providers, with community reports flagging Groq as comparatively weaker on multilingual outputs. The benchmark-visualization coding task was attempted twice per provider to assess functional code generation.

Tempr

CoverageBenchmark

Fireworks AI's August 5, 2025 technical analysis describes gpt-oss-120b and gpt-oss-20b as strong reasoning models that excel at problem solving and tool calling. Both support long context windows, configurable low/mid/high reasoning effort (mirroring o4-mini-high), and handle both built-in tools like code interpreters and user-supplied functions over multi-turn agentic trajectories. The architecture is a mixture-of-experts transformer, with performance gains attributed primarily to training data and reinforcement-learning tuning. Fireworks benchmarks place gpt-oss-120b near OpenAI o4-mini accuracy and surpassing o3-mini, while gpt-oss-20b remains surprisingly competitive despite being six times smaller. The analysis confirms OpenAI's first open-weight LLM release since GPT-2 and emphasizes suitability for agentic use cases requiring reasoning depth control and tool integration. Fireworks makes both model sizes available through its own hosted inference platform.

Tempr

CoverageBenchmark

OpenAI released gpt-oss-120b on August 5, 2025, alongside the smaller gpt-oss-20b, distributing weights, inference implementations, tool environments, and tokenizers under the Apache 2.0 license. The model is an open-weight reasoning language model trained with large-scale distillation and reinforcement learning, positioned for deep-research browsing, Python tool use, and developer-provided functions. The 120B model uses a mixture-of-experts transformer with 36 layers, alternating banded-window and fully dense attention, Grouped Query Attention with learned attention-sink biases, RoPE positional embeddings extended via YaRN, and gated SwiGLU MoE activations. More than 90% of parameters are natively quantized to MXFP4 at 4.25 bits, letting it fit on a single 80 GB GPU such as an NVIDIA H100 or AMD MI300X.

Videos about GPT OSS 120B

More models around GPT OSS 120B