Currently listed through these providers:
Model details
gpt-oss-20b
GPT-OSS 20B is the smaller sibling in OpenAI's gpt-oss family, designed for low-latency inference, local deployment, and specialized developer use cases rather than the heaviest production workloads. It uses a Mixture-of-Experts architecture with about 21B total parameters but only 3.6B active per forward pass, which keeps memory requirements around 16GB and allows it to run on consumer or single-GPU hardware. The model is released under the permissive Apache 2.0 license, is fully fine-tunable, and exposes configurable reasoning effort at low, medium, and high levels so users can trade depth for speed. Because both models in the family were trained on OpenAI's Harmony response format, it must be prompted in that format to behave correctly, and it offers full chain-of-thought access intended for debugging rather than end-user display.
The model is part of OpenAI's open-weight push to make strong reasoning and agentic capabilities widely available, with a native tool stack that covers function calling, tool use, and structured outputs, plus a 131K-token context window suitable for long-form code and document workflows. Sources describe it as matching o3-mini on common benchmarks while being small enough to fit on edge or workstation hardware, making it well suited to on-device assistants, prototyping, and customized fine-tunes. Its fine-tunability and Harmony-based training pipeline point to a forward-looking role as a flexible foundation that developers can specialize for domain-specific agents, latency-sensitive applications, and cost-conscious deployments where a larger frontier model would be overkill.
Quick Info
Powered by- Provider
- Weights & Biases
- Model key
- openai/gpt-oss-20b
- Release date
- Aug 5, 2025
- Last updated
- Aug 5, 2025
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.03
- Output token cost
- $0.13
Limits
- Output tokens
- 131,072 tokens
- Context window
- 131,072 tokens
Transparent token rates
Compare gpt-oss-20b pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about gpt-oss-20b
Videos about gpt-oss-20b
More models around gpt-oss-20b
This exact model name is also listed by 18 other providers.