Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Vertex logo

Model details

GPT OSS 20B

GPT OSS 20B is built around a compact Mixture-of-Experts architecture that uses SwiGLU activations and a token-choice routing strategy, allowing only 3.6 billion active parameters per token while maintaining strong reasoning performance. The model incorporates an alternating attention mechanism that alternates between full and sliding window contexts, paired with a learned attention sink architecture designed to optimize memory usage during inference. This architectural design was explicitly engineered for single-GPU efficiency, enabling the model to run on consumer hardware and edge devices with as little as 16 GB of memory—making advanced AI capabilities accessible beyond traditional data center infrastructure.

The model was trained using reinforcement learning techniques informed by OpenAI's most advanced internal systems, including o3 and other frontier models, and underwent comprehensive safety evaluation with global community feedback integration to build resistance against malicious fine-tuning. It operates on the Harmony response format and exposes full chain-of-thought reasoning traces, with configurable reasoning effort levels that let developers trade off latency against depth based on their specific use case. The Apache 2.0 license and full parameter fine-tuning support make it straightforward to customize for specialized applications, while its native tool use and function calling capabilities position it well for agentic workflows and production deployments that require reliable tool integration.

Vertexopenai/gpt-oss-20b-maasgpt-ossdeprecated

Quick Info

Powered by
Provider
Vertex
Model key
openai/gpt-oss-20b-maas
Release date
Aug 5, 2025
Last updated
Aug 5, 2025
AI SDK package
@ai-sdk/openai-compatible
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.07
Output token cost
$0.25

Limits

Output tokens
32,768 tokens
Context window
131,072 tokens

Transparent token rates

Compare GPT OSS 20B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GPT OSS 20B

Vertex

Coverage

According to Baidu Wiki's community-curated entry, OpenAI released gpt-oss-20b on August 5, 2025 as an open-weight AI model with 21 billion total parameters and 3.6 billion activated per token. It uses a Mixture-of-Experts Transformer architecture combining dense attention with local banded sparse attention, supports a The entry reports that Microsoft made gpt-oss-20b available to Windows 11 users via the Windows AI Foundry platform on August 6, 2025 for local invocation, and that SiliconFlow International launched the model on August 19, 2025. Citing OpenAI's technical report, the entry states gpt-oss-20b's performance is comparable

Videos about GPT OSS 20B

More models around GPT OSS 20B