Sulat.com
AI models
OpenRouter logo

Model details

GPT OSS 20B

gpt-oss-20b represents OpenAI's move into fully open-source deployment with a compact Mixture-of-Experts design that balances reasoning power against resource constraints. The model houses roughly 21 billion total parameters but activates only 3.6 billion per forward pass through its 32-expert routing system, allowing it to deliver sophisticated chain-of-thought reasoning while staying lean enough for single-GPU setups and high-end consumer laptops with 16–32 GB of RAM. Architectural choices include SwiGLU activations, an alternating attention mechanism that blends full and sliding window contexts, and a learned attention sink for memory efficiency—all decisions aimed at keeping latency low without sacrificing depth. Native FP4 quantization support further accelerates inference, making this model unusually practical for developers who want OpenAI-grade reasoning capability without data-center infrastructure.

The model's training incorporated comprehensive safety evaluation, community feedback integration, and verification against malicious fine-tuning attempts, reflecting OpenAI's effort to make openness responsible rather than reckless. It ships in the Harmony response format with the standard GPT-4o tokenizer, enabling straightforward integration into existing pipelines. gpt-oss-20b ships with configurable reasoning effort across low, medium, and high settings, letting users trade speed for depth depending on the task. Support for fine-tuning, function calling, tool use, and structured outputs positions it well for agentic applications—coding assistants, RAG workflows, and autonomous problem-solving pipelines—where the blend of open weights, efficient scaling, and agentic capability gives developers both control and power.

OpenRouteropenai/gpt-oss-20bgpt-oss

Quick Info

Powered by
Provider
OpenRouter
Model key
openai/gpt-oss-20b
Release date
Aug 5, 2025
Last updated
Aug 5, 2025
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.03
Output token cost
$0.13

Limits

Output tokens
117,964 tokens
Context window
131,072 tokens

Transparent token rates

Compare gpt-oss pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GPT OSS 20B

OpenRouter

Official sourceBenchmark

gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. $0.029 per million input tokens, $0.14 per million output tokens. 131,072 token context window. Higher uptime with 14 providers. Includes independent benchmarks from Artificial Analysis.

OpenRouter

Official sourceComparison

Compare gpt-oss-20b (free) from OpenAI to other AI models on key metrics including benchmarks, price, context length, and other model features.

OpenRouter

Official sourceBenchmark

gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. $0 per million input tokens, $0 per million output tokens. 131,072 token context window, maximum output of 8,192 tokens. Higher uptime with 14 providers. Includes independent benchmarks from Artificial Analysis.

Videos about GPT OSS 20B

More models around GPT OSS 20B