Currently listed through these providers:
Model details
Kimi K3 Fast
Kimi K3 Fast is positioned as a faster serving variant of Moonshot AI's Kimi K3, retaining the full 1,048,576-token context window while routing requests through providers such as Fireworks and Morph to favor lower latency over per-token cost. It accepts both text and image inputs and produces text outputs, with reasoning and tool-use capability surfaced by both the Fireworks AI listing and the Vercel AI Gateway documentation. Use of the model is subject to Moonshot AI's published terms and privacy policy, indicating that Moonshot AI remains the underlying rights holder even when the inference path is hosted elsewhere.
On the Fireworks AI router, the model is exposed under the identifier accounts/fireworks/routers/kimi-k3-fast, while the Vercel AI Gateway surfaces it to the AI SDK as moonshotai/kimi-k3-fast, giving developers a stable name to reference across stack integrations. The open-weights designation, combined with tool calling and structured output support, makes the model a practical fit for agent-style workflows that need long-context retrieval over documents and images and want to keep response latency tight for interactive applications.
Quick Info
Powered by- Provider
- Neuralwatt
- Model key
- kimi-k3-fast
- Release date
- Jul 16, 2026
- Last updated
- Jul 16, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $3.00
- Output token cost
- $15.00
Limits
- Output tokens
- 1,048,560 tokens
- Context window
- 1,048,560 tokens
Transparent token rates
Compare Kimi K3 Fast pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about Kimi K3 Fast
No articles yet. Fetch the latest news to show it here.
Videos about Kimi K3 Fast
More models around Kimi K3 Fast
This exact model name is also listed by 4 other providers.