FastRouter
Major open-source foundational model releases are an exciting time for the AI community, bringing unique architectural innovations and capabilities.
Model details
GPT OSS 20B is part of OpenAI's gpt-oss family of open-weight large language models, released as Apache 2.0 licensed weights that can be downloaded, run locally, and fine-tuned on user-controlled hardware. The family marked OpenAI's first public release of internal model parameters since GPT-2, framing the gpt-oss lineup as an alternative to API-locked proprietary systems. Community coverage of the announcement appeared on the same date as the model's availability, positioning the 20B variant alongside its larger sibling as a flexible option for developers who want transparency and on-device control.
Beyond the model weights themselves, an MXFP4-quantized GGUF build of GPT-OSS-20B is distributed through community channels for local inference, with the Qualcomm AI Hub listing a roughly 11.27 GB artifact packaged for Snapdragon-class devices. The same deployment listing links the model to arXiv paper 2508.10925 as the associated research reference and points to a Hugging Face repository at qualcomm/GPT-OSS-20B where the quantized files are hosted. A llama.cpp-based runtime path is documented for Windows, Android, and Linux, with reported prefill and decode throughput on Snapdragon X2 Elite hardware, illustrating the model's fit for edge and local experimentation rather than purely cloud-hosted serving.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
FastRouter
Major open-source foundational model releases are an exciting time for the AI community, bringing unique architectural innovations and capabilities.
Opper
OpenAI released gpt-oss-20b on August 5, 2025 as an open-weight model with 21 billion total parameters, of which 3.6 billion activate per token, built around a Mixture-of-Experts Transformer architecture that combines dense attention with a local banded sparse attention mechanism and supports a 128,000-token context le According to the model's technical report, its performance is comparable to o3-mini, with strong results on benchmarks such as AIME and HealthBench, and in local deployment tests a quantized GPT-OSS 20B weighing 11.95 GB loaded in 8.3 seconds. The encyclopedia entry also documents distribution milestones: on August 6,
This exact model name is also listed by 20 other providers.