GPT OSS 20B sits at the lighter end of OpenAI's gpt-oss family, released as an open-weight model under the permissive Apache 2.0 license alongside its larger sibling. The model carries 21B total parameters with 3.6B active parameters, a sparse configuration that lets it target lower-latency, on-device, or specialized deployments where running a frontier-scale model is impractical. Like its sibling, it was trained specifically on OpenAI's harmony response format, so downstream systems must wrap prompts and completions in that format to get correct behavior; using a generic chat template will degrade output quality.
For practical work, the model is designed to expose its reasoning process end-to-end, with configurable reasoning effort across low, medium, and high settings so developers can trade latency against depth of thought on a per-request basis. It supports parameter fine-tuning and ships with native function-calling for agentic pipelines, making it a flexible base for building custom assistants, retrieval-augmented systems, or domain-specific copilots that need to run locally. The combination of permissive licensing, adjustable reasoning, and an accessible parameter footprint positions it well for experimentation, private deployments, and cost-sensitive production use where full-size frontier models would be overkill.