Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Tempr Gateway logo

Model details

GPT OSS 120B

GPT OSS 120B is OpenAI's open-weight language model built around a Mixture-of-Experts architecture with 117B total parameters, activating 5.1B parameters per forward pass across 128 experts. This sparse-activation design lets a very large model behave like a much smaller one at inference time, trading raw scale for efficient compute per token. The model is positioned for high-reasoning, agentic, and general-purpose production use cases, offering configurable reasoning depth and full chain-of-thought access so developers can tune how much deliberation the model performs versus how quickly it responds. Native tool use is built in, including function calling, web browsing, and structured output generation, which makes it well suited to multi-step workflows where the model needs to call external APIs, retrieve information, and return machine-readable results.</placeholder>

Weights for GPT OSS 120B are openly published, allowing teams to self-host, fine-tune, or audit the model rather than depending solely on a hosted API. Its long context window and text-in, text-out design fit production pipelines that combine reasoning with tool orchestration, such as research assistants, coding agents, and automated analytics. The sparse MoE structure is aimed at practical deployment efficiency rather than maximum parameter count on a single forward pass, making the model attractive for organizations that want frontier-style reasoning behavior with the cost profile of a smaller active model. Developers integrating GPT OSS 120B typically pair it with structured-output schemas and external tools to take advantage of its chain-of-thought transparency and native function-calling support.

Tempr Gatewaycerebras/gpt-oss-120bgpt-oss

Quick Info

Powered by
Provider
Tempr Gateway
Model key
cerebras/gpt-oss-120b
Release date
Aug 5, 2025
Last updated
Jun 10, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.35
Output token cost
$0.75

Limits

Output tokens
40,960 tokens
Context window
131,072 tokens

Transparent token rates

Compare GPT OSS 120B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GPT OSS 120B

Impossibl

Coverage

OpenAI released gpt-oss-120b and gpt-oss-20b as open-weight language models on August 5, 2025, its first open-weight release since GPT-2 in 2019. The text-only models target lower-cost, customizable deployments for developers, researchers, and enterprises, distinguishing open-weight from open-source licensing. OpenAI President Greg Brockman described the models as a contribution to the open-model ecosystem alongside releases from Meta, Mistral, and DeepSeek. OpenAI confirmed partnerships with Nvidia, AMD, Cerebras, and Groq to ensure the models run efficiently across hardware platforms, with extensive safety training completed prior to release.

Impossibl

CoverageBenchmark

Fireworks AI published a technical deep-dive on OpenAI's gpt-oss-120b (and 20b), released August 5, 2025. The analysis covers the model's standard mixture-of-experts transformer architecture, reinforcement-learning training improvements, and benchmark positioning at parity with OpenAI's o3 and o4-mini at high reasoning levels. The piece frames gpt-oss-120b as a strong reasoning model well suited to agentic workflows. Both gpt-oss variants support configurable low/mid/high reasoning effort, long context windows, and tool use, including built-in code interpreter and browser tools plus user-provided function calls over multi-turn trajectories. Fireworks also notes OpenAI collaborated with Nvidia, AMD, Cerebras, and Groq to ensure broad hardware compatibility for the open-weight release.

Impossibl

CoverageBenchmark

OpenAI's gpt-oss-120b is an open-weight reasoning model released on August 5, 2025 under the Apache 2.0 license, alongside the smaller gpt-oss-20b. Per the model card and Hugging Face listing, it has 117B total parameters with 5.1B active per token and a 131K-token context window, fitting on a single 80 GB GPU such as an NVIDIA H100 or AMD MI300X. The architecture is an autoregressive Mixture-of-Experts transformer with 36 layers alternating between banded sliding-window attention and fully dense attention, plus Grouped Query Attention with learned attention-sink biases, RoPE positional embeddings extended via YaRN, gated SwiGLU MoE activations, and native MXFP4 quantization at roughly 4.25 bits per parameter for the MoE weights.

Videos about GPT OSS 120B

More models around GPT OSS 120B