Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
DevPass (LLM Gateway) logo

Model details

GPT OSS 20B

GPT OSS 20B is part of OpenAI's gpt-oss family of open-weight large language models, released as Apache 2.0 licensed weights that can be downloaded, run locally, and fine-tuned on user-controlled hardware. The family marked OpenAI's first public release of internal model parameters since GPT-2, framing the gpt-oss lineup as an alternative to API-locked proprietary systems. Community coverage of the announcement appeared on the same date as the model's availability, positioning the 20B variant alongside its larger sibling as a flexible option for developers who want transparency and on-device control.

Beyond the model weights themselves, an MXFP4-quantized GGUF build of GPT-OSS-20B is distributed through community channels for local inference, with the Qualcomm AI Hub listing a roughly 11.27 GB artifact packaged for Snapdragon-class devices. The same deployment listing links the model to arXiv paper 2508.10925 as the associated research reference and points to a Hugging Face repository at qualcomm/GPT-OSS-20B where the quantized files are hosted. A llama.cpp-based runtime path is documented for Windows, Android, and Linux, with reported prefill and decode throughput on Snapdragon X2 Elite hardware, illustrating the model's fit for edge and local experimentation rather than purely cloud-hosted serving.

DevPass (LLM Gateway)gpt-oss-20bgpt-oss

Quick Info

Powered by
Provider
DevPass (LLM Gateway)
Model key
gpt-oss-20b
Release date
Aug 5, 2025
Last updated
Aug 5, 2025
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.04
Output token cost
$0.19

Limits

Output tokens
32,766 tokens
Context window
131,072 tokens

Transparent token rates

Compare GPT OSS 20B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GPT OSS 20B

LLM Gateway

CoverageComparison

Comparison: Qwen3:30b vs GPT-OSS:20b Here is a comparison between Qwen3:30b and GPT-OSS:20b focusing on instruction following and performance parameters, specs and speed: Architecture and …

Opper

Coverage

OpenAI released gpt-oss-20b on August 5, 2025 as an open-weight model with 21 billion total parameters, of which 3.6 billion activate per token, built around a Mixture-of-Experts Transformer architecture that combines dense attention with a local banded sparse attention mechanism and supports a 128,000-token context le According to the model's technical report, its performance is comparable to o3-mini, with strong results on benchmarks such as AIME and HealthBench, and in local deployment tests a quantized GPT-OSS 20B weighing 11.95 GB loaded in 8.3 seconds. The encyclopedia entry also documents distribution milestones: on August 6,

LLM Gateway

Official sourceBenchmark

LLM Gateway lists GPT OSS 20B as a STABLE model with a 131,072-token context window and capability flags for Streaming, Tools, Reasoning, and JSON Output, routed via an OpenAI-compatible endpoint at https://api.llmgateway.io/v1. The gateway currently exposes three upstream providers for this model: Groq ($0.10/M input, Reported 24-hour performance on the gateway is thin: Groq shows only 15 requests and NanoGPT only 3, so throughput, uptime, and error metrics are suppressed and cannot be used to compare providers reliably. The page has no benchmarks, no independent quality scores, and no release date, so developers evaluating GPT OSS

Videos about GPT OSS 20B

More models around GPT OSS 20B