GPT OSS 120B is OpenAI's open-weight language model built around a Mixture-of-Experts architecture with 117B total parameters, activating 5.1B parameters per forward pass across 128 experts. This sparse-activation design lets a very large model behave like a much smaller one at inference time, trading raw scale for efficient compute per token. The model is positioned for high-reasoning, agentic, and general-purpose production use cases, offering configurable reasoning depth and full chain-of-thought access so developers can tune how much deliberation the model performs versus how quickly it responds. Native tool use is built in, including function calling, web browsing, and structured output generation, which makes it well suited to multi-step workflows where the model needs to call external APIs, retrieve information, and return machine-readable results.</placeholder>
Weights for GPT OSS 120B are openly published, allowing teams to self-host, fine-tune, or audit the model rather than depending solely on a hosted API. Its long context window and text-in, text-out design fit production pipelines that combine reasoning with tool orchestration, such as research assistants, coding agents, and automated analytics. The sparse MoE structure is aimed at practical deployment efficiency rather than maximum parameter count on a single forward pass, making the model attractive for organizations that want frontier-style reasoning behavior with the cost profile of a smaller active model. Developers integrating GPT OSS 120B typically pair it with structured-output schemas and external tools to take advantage of its chain-of-thought transparency and native function-calling support.