Currently listed through these providers:
Model details
GPT-5.6 Luna (US)
GPT-5.6 Luna is presented as a fast, cost-efficient model aimed at high-volume tasks such as classification, extraction, summarization, and lightweight agent steps. According to its gateway description, the model is engineered for production pipelines where throughput and predictable cost matter more than deep reasoning, with text and image inputs supported over a very large that quick-info value window in the million-token range. That positioning frames it as a workhorse layer in agent stacks, where it can handle triage, routing, and short-form generation before heavier models are called in for complex synthesis.
In practice, the model is exposed through a managed Bedrock deployment identifier and supports server-side tool calling and prompt caching, which makes it well suited to long, repeated contexts that benefit from cache reuse. The deployment reports a roughly 1.1M-token that quick-info value window with around the cataloged API limit tokens of output headroom, and a gateway-listed capability set that includes vision, reasoning, tool calling, caching, web search, and JSON schema support. Because most evidence comes from a third-party router page rather than first-party model documentation, technical details such as training data composition, parameter count, and benchmark scores are not verifiable here, so the practical takeaway is that GPT-5.6 Luna is best understood as a lean, context-rich OpenAI model for classification, extraction, summarization, and lightweight agent steps on Bedrock.
Quick Info
Powered by- Provider
- Amazon Bedrock
- Model key
- us.openai.gpt-5.6-luna
- Release date
- Jul 9, 2026
- Last updated
- Jul 9, 2026
- Knowledge cutoff
- 2026-02-16
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.22
- Output token cost
- $1.32
Limits
- Input tokens
- 922,000 tokens
- Output tokens
- 128,000 tokens
- Context window
- 1,050,000 tokens
Transparent token rates
Compare GPT-5.6 Luna (US) pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about GPT-5.6 Luna (US)
No articles yet. Fetch the latest news to show it here.
Videos about GPT-5.6 Luna (US)
More models around GPT-5.6 Luna (US)
This exact model name is also listed by 34 other providers.