Currently listed through these providers:
Model details
OpenAI GPT 120B
OpenAI GPT 120B is a text-to-text reasoning model served through an OpenAI-compatible completions endpoint at inference.baseten.co/v1, where Baseten's reasoning documentation explicitly enables extended thinking by default for the openai/gpt-oss-120b slug. The deployment exposes a generous 128,072-token context window paired with an identical 128,072-token output cap, giving applications room to hold long documents, multi-turn agent traces, or tool outputs in working memory while still producing lengthy completions. As part of the gpt-oss family of open-weight models, it is well suited to teams that want a capable reasoning model they can run on hosted infrastructure without standing up their own serving stack, especially for workflows that benefit from chain-of-thought traces surfaced as separate reasoning content alongside the final answer.
Practical usage benefits from the model's OpenAI-style reasoning effort control, which maps the familiar off, minimal, low, medium, high, xhigh, and max levels into OpenAI's thinking format and is supported in streaming responses, making it straightforward to tune latency versus answer quality for production agents. Pricing is positioned at $0.10 per million input tokens and $0.50 per million output tokens in USD, with cache read and cache write rates listed at $0, which keeps cost predictable for cached, retrieval-heavy pipelines. Because the endpoint follows the openai-completions API contract and supports strict mode for structured output, developers can integrate it into existing OpenAI-style toolchains, function-calling frameworks, and agent runtimes with minimal changes, leaning on Baseten's hosted inference while retaining access to chain-of-thought reasoning when needed.
Quick Info
Powered by- Provider
- Baseten
- Model key
- openai/gpt-oss-120b
- Release date
- Aug 5, 2025
- Last updated
- Aug 5, 2025
- Knowledge cutoff
- 2025-08
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.10
- Output token cost
- $0.50
Limits
- Output tokens
- 128,072 tokens
- Context window
- 128,072 tokens
Transparent token rates
Compare OpenAI GPT 120B pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about OpenAI GPT 120B
No articles yet. Fetch the latest news to show it here.