Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Qiniu logo

Model details

gpt-oss-20b

gpt-oss-20b is the smaller member of OpenAI's open-weight gpt-oss series, designed for lower-latency, local, or specialized deployments rather than the largest production workloads. The model card describes it as carrying 21B total parameters with 3.6B active parameters, giving it a comparatively light inference footprint relative to its 117B/5.1B sibling. It is published under the permissive Apache 2.0 license, and OpenAI frames the series as built for powerful reasoning, agentic tasks, and versatile developer use cases, with full chain-of-thought exposed for debugging and configurable low, medium, and high reasoning effort to balance latency against quality.

Because the model is trained on OpenAI's harmony response format and is meant to be used with that format, it slots naturally into agentic pipelines that need function calling and structured tool use, while still allowing parameter fine-tuning for narrower domains. Its smaller active-parameter budget makes it well suited to single-GPU or local setups where the 120B variant would be impractical, and community work like the GPT-OSS-20B-Vision project shows the open weights can be extended into multimodal territory with relatively modest hardware. For teams wanting an open, reasoning-oriented model that runs outside a hyperscale cluster, gpt-oss-20b offers a pragmatic middle ground between local models and frontier-scale deployments.

Qiniugpt-oss-20b

Quick Info

Powered by
Provider
Qiniu
Model key
gpt-oss-20b
Release date
Aug 6, 2025
Last updated
Aug 6, 2025
Input modalities
Output modalities
Capabilities

Limits

Output tokens
4,096 tokens
Context window
128,000 tokens

Latest news about gpt-oss-20b

OVHcloud AI Endpoints

Coverage

A cross-vendor open-weights LLM release history and timeline published on hidekazu-konishi.com covers OpenAI's gpt-oss family alongside Llama, Mistral, Qwen, DeepSeek, and milestone references to Gemma and Phi. The article explicitly names gpt-oss-20b as part of the gpt-oss family and links release rows to official fir First published July 26, 2026 and last updated August 24, 2026, the timeline provides structured reference context for gpt-oss-20b's place in the broader open-weights landscape, useful for developers comparing model generations and licensing terms. The entry does not address OVHcloud-specific hosting or serving changes

Regolo AI

Coverage

The Baidu encyclopedia page on GPT-OSS-20B records that OpenAI released this open-weight model on August 5, 2025, with 21 billion total parameters and 3.6 billion active parameters per token, designed for low-latency local use on edge devices with 16GB of memory. It documents the model's Mixture-of-Experts Transformer The same page notes that GPT-OSS-20B's performance is reportedly comparable to OpenAI's o3-mini, with strong results on AIME and HealthBench, while some evaluations indicate its overall SWE-bench Verified performance is surpassed by Zhipu GLM-4.7-Flash. It also reports a 2026 study by institutions including The Univers

OVHcloud AI Endpoints

CoverageAnalysis

arXiv preprint 2508.16700 presents a single-GPU deployment evaluation of GPT-OSS-20B on NVIDIA H100 hardware in bf16 precision, explicitly naming gpt-oss-20b and characterizing it as a Mixture-of-Experts model with 20.9 billion total parameters and approximately 3.61 billion active parameters (about 17.3%). The study m At a 2048-token context with 64-token decode, the paper reports gpt-oss-20B delivers about 31.8% higher decode throughput, 25.8% lower energy per 1000 generated tokens, and 31.7% less peak VRAM than Qwen3-32B, while noting higher TTFT due to MoE routing overhead. It introduces an Active Parameter Efficiency (APE) metri

Videos about gpt-oss-20b