Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Baseten logo

Model details

OpenAI GPT 120B

OpenAI GPT 120B is a text-to-text reasoning model served through an OpenAI-compatible completions endpoint at inference.baseten.co/v1, where Baseten's reasoning documentation explicitly enables extended thinking by default for the openai/gpt-oss-120b slug. The deployment exposes a generous 128,072-token context window paired with an identical 128,072-token output cap, giving applications room to hold long documents, multi-turn agent traces, or tool outputs in working memory while still producing lengthy completions. As part of the gpt-oss family of open-weight models, it is well suited to teams that want a capable reasoning model they can run on hosted infrastructure without standing up their own serving stack, especially for workflows that benefit from chain-of-thought traces surfaced as separate reasoning content alongside the final answer.

Practical usage benefits from the model's OpenAI-style reasoning effort control, which maps the familiar off, minimal, low, medium, high, xhigh, and max levels into OpenAI's thinking format and is supported in streaming responses, making it straightforward to tune latency versus answer quality for production agents. Pricing is positioned at $0.10 per million input tokens and $0.50 per million output tokens in USD, with cache read and cache write rates listed at $0, which keeps cost predictable for cached, retrieval-heavy pipelines. Because the endpoint follows the openai-completions API contract and supports strict mode for structured output, developers can integrate it into existing OpenAI-style toolchains, function-calling frameworks, and agent runtimes with minimal changes, leaning on Baseten's hosted inference while retaining access to chain-of-thought reasoning when needed.

Basetenopenai/gpt-oss-120bgpt-oss

Quick Info

Powered by
Provider
Baseten
Model key
openai/gpt-oss-120b
Release date
Aug 5, 2025
Last updated
Aug 5, 2025
Knowledge cutoff
2025-08
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.10
Output token cost
$0.50

Limits

Output tokens
128,072 tokens
Context window
128,072 tokens

Transparent token rates

Compare OpenAI GPT 120B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about OpenAI GPT 120B

No articles yet. Fetch the latest news to show it here.

Videos about OpenAI GPT 120B

More models around OpenAI GPT 120B