Sulat.com
AI models
Weights & Biases logo

Model details

gpt-oss-20b

GPT-OSS 20B is the smaller sibling in OpenAI's gpt-oss family, designed for low-latency inference, local deployment, and specialized developer use cases rather than the heaviest production workloads. It uses a Mixture-of-Experts architecture with about 21B total parameters but only 3.6B active per forward pass, which keeps memory requirements around 16GB and allows it to run on consumer or single-GPU hardware. The model is released under the permissive Apache 2.0 license, is fully fine-tunable, and exposes configurable reasoning effort at low, medium, and high levels so users can trade depth for speed. Because both models in the family were trained on OpenAI's Harmony response format, it must be prompted in that format to behave correctly, and it offers full chain-of-thought access intended for debugging rather than end-user display.

The model is part of OpenAI's open-weight push to make strong reasoning and agentic capabilities widely available, with a native tool stack that covers function calling, tool use, and structured outputs, plus a 131K-token context window suitable for long-form code and document workflows. Sources describe it as matching o3-mini on common benchmarks while being small enough to fit on edge or workstation hardware, making it well suited to on-device assistants, prototyping, and customized fine-tunes. Its fine-tunability and Harmony-based training pipeline point to a forward-looking role as a flexible foundation that developers can specialize for domain-specific agents, latency-sensitive applications, and cost-conscious deployments where a larger frontier model would be overkill.

Weights & Biasesopenai/gpt-oss-20bgpt-oss

Quick Info

Powered by
Provider
Weights & Biases
Model key
openai/gpt-oss-20b
Release date
Aug 5, 2025
Last updated
Aug 5, 2025
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.03
Output token cost
$0.13

Limits

Output tokens
131,072 tokens
Context window
131,072 tokens

Transparent token rates

Compare gpt-oss-20b pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about gpt-oss-20b

Videos about gpt-oss-20b

More models around gpt-oss-20b