Model details
gpt-oss-120b-fast
The gpt-oss-120b-fast is a Mixture of Experts model built by OpenAI as part of their open-weight initiative, designed to deliver powerful reasoning and agentic capabilities in a production-friendly package. Its architecture supports configurable reasoning effort, allowing developers to tune the model's thinking depth between low, medium, and high settings depending on the task's complexity and latency requirements. The model provides full chain-of-thought access, giving visibility into the internal reasoning process for easier debugging and trust-building. MXFP4 quantization enables deployment on a single 80GB GPU such as an NVIDIA H100 or AMD MI300X, dramatically reducing hardware barriers for teams that want to run large-scale reasoning models without distributed infrastructure.
The model was trained specifically on OpenAI's harmony response format, which structures how prompts and outputs are exchanged. This format integration means the model expects inputs formatted in a particular way to function correctly, and it excels in developer workflows that align with this structure. The Apache 2.0 license removes copyleft restrictions and patent risk, making it suitable for commercial products, experimentation, and customization without legal friction. With its open-weight status, reasoning-first design, and efficient quantization pathway, this model targets teams building agents, automated workflows, and applications that benefit from transparent, inspectable reasoning chains in production environments.
Quick Info
Powered by- Provider
- Nebius Token Factory
- Model key
- openai/gpt-oss-120b-fast
- Release date
- Jun 10, 2025
- Last updated
- May 7, 2026
- Knowledge cutoff
- 2025-06
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.10
- Output token cost
- $0.50
Limits
- Input tokens
- 7,000 tokens
- Output tokens
- 8,192 tokens
- Context window
- 8,000 tokens