Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
LLM Gateway logo

Model details

Kimi K2.7 Code Highspeed (Tencent Cloud)

Kimi K2.7 Code HighSpeed is presented as a serving-layer variant of Moonshot AI's K2.7 Code coding model, sharing the same trillion-parameter MoE architecture and MoonViT 400M vision encoder while pushing output throughput higher for real-time agentic coding. The HighSpeed profile is described as reaching roughly 260 tokens per second at peak on short-context tasks, about six times the speed of the standard K2.7 Code serving mode, which makes it well suited to workflows where an AI agent must generate many small pieces of code in rapid succession. It also reportedly uses about 30% fewer reasoning tokens than the prior K2.6 generation, helping keep latency low during iterative tool calls and code edits without changing the underlying model behavior.

The variant retains K2.7 Code's mandatory thinking mode and is positioned for developer and team use cases such as in-editor code completion, multi-step refactoring agents, and high-volume batch generation through Kimi Code Beta, API, and Business tiers. It carries a 256K context window and is distributed as open weights on HuggingFace under a Modified MIT License, supporting self-hosting and integration into custom coding pipelines. The HighSpeed optimization is framed as an infrastructure-level change rather than a new model, so teams adopting it gain throughput for agentic loops while keeping the coding quality profile established by the K2.7 Code release.

LLM Gatewaytencent/kimi-k2.7-code-highspeedkimi-k2

Quick Info

Powered by
Provider
LLM Gateway
Model key
tencent/kimi-k2.7-code-highspeed
Release date
Jun 12, 2026
Last updated
Jun 12, 2026
Knowledge cutoff
2025-01
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$1.90
Output token cost
$8.00

Limits

Output tokens
262,144 tokens
Context window
262,144 tokens

Transparent token rates

Compare Kimi K2.7 Code Highspeed (Tencent Cloud) pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Kimi K2.7 Code Highspeed (Tencent Cloud)

LLM Gateway

CoverageBenchmark

Moonshot AI released Kimi K2.7-Code on June 12, 2026, and on June 15, 2026 announced a HighSpeed Mode variant rolling out to Kimi Code Beta, according to TechTimes reporting on the same-day Moonshot post. The HighSpeed variant reportedly delivers around 180 tokens per second on median-length coding inputs and up to 260 The article notes that Moonshot has not submitted K2.7-Code to any independent coding benchmark, so all efficiency claims — including the roughly 30% reduction in reasoning-token usage versus K2.6 and the new HighSpeed throughput figures — rest on Moonshot's own five proprietary benchmarks. K2.7-Code is framed as the f

Videos about Kimi K2.7 Code Highspeed (Tencent Cloud)

More models around Kimi K2.7 Code Highspeed (Tencent Cloud)