Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Hugging Face logo

Model details

Kimi-K2.5

Kimi-K2.5 is a multimodal agentic model engineered to advance general agentic intelligence by seamlessly integrating vision and language processing. Built upon a trillion-parameter mixture-of-experts transformer architecture, the model is designed to treat text and visual inputs as mutually enhancing components rather than competing data streams. Its primary innovation is the Agent Swarm framework, which allows the model to dynamically decompose complex, heterogeneous problems into smaller sub-tasks. By executing these tasks concurrently, the architecture significantly reduces latency compared to traditional single-agent approaches, making it highly effective for sophisticated coding, reasoning, and computer-use applications.

The model lineage stems from continual pre-training on approximately 15 trillion tokens, which establishes a robust foundation for its specialized operational modes. Through this training, Kimi-K2.5 achieves strong performance on benchmarks like Humanity’s Last Exam, demonstrating high efficiency in both execution speed and cost. Its design supports a variety of paradigms, including instant and thinking modes, which allow users to tailor the model to specific conversational or agentic requirements. By providing this capability as an open-weight model, it serves as a versatile tool for researchers and developers looking to deploy scalable, parallelized agentic systems in real-world production environments.

Hugging Facemoonshotai/Kimi-K2.5kimi-k2

Quick Info

Powered by
Provider
Hugging Face
Model key
moonshotai/Kimi-K2.5
Release date
Jan 1, 2026
Last updated
Jan 1, 2026
Knowledge cutoff
2025-01
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.60
Output token cost
$3.00

Limits

Output tokens
262,144 tokens
Context window
262,144 tokens

Transparent token rates

Compare Kimi-K2.5 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Kimi-K2.5

Hugging Face

Official sourceAnnouncement

A Blog post by Leco Li on Hugging Face

Tencent Coding Plan (China)

CoverageBenchmark

A Lorphic explainer dated July 13, 2026 walks through the Kimi K2 model family (K2, K2.5, K2.6, K2.7), emphasizing that the variants are architecturally related but carry different capabilities, licensing implications, and recommended use cases. It attributes to Moonshot's technical documentation and Hugging Face model The same explainer states the K2 base model was pre-trained on 15.5 trillion tokens and that Moonshot used the Muon optimizer for the family, framing these as the architectural foundation that all K2 variants (including K2.5) share. It positions the article as a practical breakdown covering architecture, version-by-ver

Videos about Kimi-K2.5

More models around Kimi-K2.5