gpt-oss-120b is OpenAI's flagship open-weight reasoning model, released under Apache 2.0 as part of the gpt-oss family and intended for high-reasoning, agentic, and general-purpose production workloads. It is built as a 117-billion-parameter Mixture-of-Experts transformer that activates only about 5.1 billion parameters per forward pass, with native MXFP4 quantization that lets the full model run on a single H100-class accelerator. The architecture pairs this sparse MoE design with a 131K-token context window, configurable reasoning effort (low, medium, high), and full chain-of-thought access, so the model can be tuned toward either fast throughput or deeper deliberation depending on the task.
In practical terms, gpt-oss-120b is designed to behave like a production-grade open model rather than a research artifact: it supports native function calling, browsing, and structured output, making it suitable for tool-using agents and pipelines that previously required closed APIs. Ollama Cloud exposes the same weights through its cloud-models mechanism, which offloads execution to managed infrastructure while keeping the familiar local CLI and library interface, so users without a high-end GPU can still experiment with or deploy the 120B variant. The combination of open licensing, efficient sparse activation, long context, and tool-use support makes it a strong fit for teams that want OpenAI-style reasoning with the ability to self-host, fine-tune, or route across providers.