Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Ollama Cloud logo

Model details

gpt-oss:120b

gpt-oss-120b is OpenAI's flagship open-weight reasoning model, released under Apache 2.0 as part of the gpt-oss family and intended for high-reasoning, agentic, and general-purpose production workloads. It is built as a 117-billion-parameter Mixture-of-Experts transformer that activates only about 5.1 billion parameters per forward pass, with native MXFP4 quantization that lets the full model run on a single H100-class accelerator. The architecture pairs this sparse MoE design with a 131K-token context window, configurable reasoning effort (low, medium, high), and full chain-of-thought access, so the model can be tuned toward either fast throughput or deeper deliberation depending on the task.

In practical terms, gpt-oss-120b is designed to behave like a production-grade open model rather than a research artifact: it supports native function calling, browsing, and structured output, making it suitable for tool-using agents and pipelines that previously required closed APIs. Ollama Cloud exposes the same weights through its cloud-models mechanism, which offloads execution to managed infrastructure while keeping the familiar local CLI and library interface, so users without a high-end GPU can still experiment with or deploy the 120B variant. The combination of open licensing, efficient sparse activation, long context, and tool-use support makes it a strong fit for teams that want OpenAI-style reasoning with the ability to self-host, fine-tune, or route across providers.

Ollama Cloudgpt-oss:120bgpt-oss

Quick Info

Powered by
Provider
Ollama Cloud
Model key
gpt-oss:120b
Release date
Aug 5, 2025
Last updated
Jan 19, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.15
Output token cost
$0.60

Limits

Output tokens
32,768 tokens
Context window
131,072 tokens

Transparent token rates

Compare gpt-oss:120b pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about gpt-oss:120b

Ollama Cloud

CoverageBenchmark

Core42's August 28, 2025 press release announces the global availability of OpenAI's gpt-oss-120B on Core42 AI Cloud via the Compass API, powered by Cerebras Inference at approximately 3,000 tokens per second using the Cerebras CS-3 wafer-scale engine. The release includes a quote from Trevor Cai, OpenAI's Head of Infr Cerebras CEO Andrew Feldman frames the deployment as extending strategic partnership access to enterprises, researchers, and governments in the Middle East and globally for reasoning-capable workloads. While the hosting stack (Cerebras CS-3 via Core42 Compass API) is unrelated to Ollama Cloud, the announcement demonstr

Videos about gpt-oss:120b

More models around gpt-oss:120b