Currently listed through these providers:
Model details
GPT-5.6 Luna
GPT-5.6 Luna sits inside OpenAI's GPT-5.6 series and is positioned as the speed- and cost-optimized tier of that family, designed for high-volume, latency-sensitive workloads rather than maximum-reasoning flagship use. The available listing describes it as suited to chat, classification, and lightweight agentic workflows, delivering capable reasoning relative to its price tier. That combination of a one-million-token context window with low per-token pricing makes it a practical fit for applications that need long context with high throughput, such as conversational assistants, document triage pipelines, and routine multi-step tool use where every request must stay responsive at a tight unit economics.
From a deployment standpoint, the same model is exposed through multiple routed providers, giving teams flexibility to trade off latency, uptime, and tool-calling accuracy depending on their workload. The listed standard routing shows an OpenAI-hosted instance alongside Amazon Bedrock and Azure regional options, each reporting different latency and throughput characteristics. This multi-provider footprint, combined with the model's emphasis on fast, lightweight reasoning, makes GPT-5.6 Luna a sensible default when developers want a single model that can serve both interactive chat and background classification at scale without paying flagship prices.
Quick Info
Powered by- Provider
- Kilo Gateway
- Model key
- openai/gpt-5.6-luna
- Release date
- Jul 9, 2026
- Last updated
- Jul 9, 2026
- Knowledge cutoff
- 2026-02-16
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.20
- Output token cost
- $1.20
Limits
- Input tokens
- 922,000 tokens
- Output tokens
- 128,000 tokens
- Context window
- 1,050,000 tokens
Latest news about GPT-5.6 Luna
Videos about GPT-5.6 Luna
Recent tweets and retweets from Kilo Gateway
More models around GPT-5.6 Luna
This exact model name is also listed by 35 other providers.