SiliconFlow
Built on Kimi K2 through continued pretraining over approximately 15T mixed visual and text tokens, it delivers state-of-the-art coding and vision capabilities as a native multimodal model.
Model details
Kimi K2.5 is an open-weight continuation of the Kimi K2 lineage, produced through further pretraining on roughly 15 trillion mixed visual and text tokens. That training recipe lets it function as a native multimodal model that accepts both text and images while outputting text, and SiliconFlow's launch coverage frames it as state of the art for visual agentic intelligence and coding. The continued-pretraining approach, rather than a from-scratch build, positions K2.5 as an evolution that retains K2's foundations while broadening modality coverage and agent-oriented behavior.
In practice, K2.5 is aimed at agent-style workloads: long-context multi-turn reasoning, tool calling, structured outputs, and visual understanding of interfaces, diagrams, and screenshots that an agent has to act on. Its reach extends beyond a single host, appearing on Cloudflare Workers AI alongside primitives such as Durable Objects, Workflows, and the Agents SDK for orchestrating agent systems, and on Amazon Bedrock as a managed model option. SiliconFlow also offers a direct head-to-head comparison against the original Kimi K2 Instruct, making it easier to weigh the visual and agentic gains against cost. The combination of open weights, a very large context window, and multi-host availability makes K2.5 a practical pick for teams building autonomous agents that need to see, reason, and act across long workflows.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
SiliconFlow
Built on Kimi K2 through continued pretraining over approximately 15T mixed visual and text tokens, it delivers state-of-the-art coding and vision capabilities as a native multimodal model.
SiliconFlow
Compare Kimi-K2-Instruct and Kimi-K2.5 across performance, cost, capabilities, and real-world use cases. See which model fits your needs.
SiliconFlow
SiliconFlow's official release notes document a critical operational change effective June 11, 2026: all traffic to moonshotai/Kimi-K2.5 is being automatically routed to the successor model moonshotai/Kimi-K2.6, at the same pricing with no cost change to customers. The original Kimi-K2.5 model is flagged for future dep This release-notes entry is the most developer-impacting finding for current SiliconFlow users of Kimi-K2.5, as it changes the model from a forward-looking release to a routing target being phased out in favor of Kimi-K2.6. The same notes page documents parallel deprecations affecting zai-org/GLM-4.7, zai-org/GLM-5 (→
SiliconFlow
The companion OpenRouter listing page for moonshotai/Kimi-K2.5 adds pricing-history and token-share analytics, with the listed price at $0.45/$2.25 per 1M tokens across SiliconFlow and DeepInfra. Weighted-average effective prices paid by customers are $0.2218/M input and $2.541/M output, substantially below list owing The price-history chart spans June through early September 2026, showing relative stability of listed input prices around $0.45/M. OpenRouter exposes per-provider effective input prices, cache-hit rates, and token-share breakdowns, making it the most detailed public source for comparing SiliconFlow's Kimi-K2.5 hosting