Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Moonshot AI logo

Model details

Kimi K2.7 Code HighSpeed

Kimi K2.7 Code HighSpeed is the accelerated tier of Moonshot AI's coding-focused line, sharing the same underlying weights as the standard K2.7 Code release but served with higher generation throughput. In external benchmark evaluations reported on the Kimi API platform, the K2.7 Code family significantly improves instruction compliance and long-horizon coding performance compared to K2.6, while reducing overthinking tendencies by about 30 percent on average, making it well suited for multi-step software engineering workflows that require sustained adherence to detailed prompts.

The practical advantage of this HighSpeed variant is generation rate, delivering roughly 180 tokens per second during typical coding workloads and peaks near 260 tokens per second in shorter context scenarios, which translates into noticeably snappier autocomplete, refactoring, and interactive debugging sessions. It accepts text and image inputs while emitting text and tool-use outputs, and supports tool calling, tool choice, structured output, and streaming, giving agentic code assistants and IDE plug-ins a responsive interface to functions, schemas, and multi-turn dialog without giving up the model's long-context coding strengths.

Moonshot AIkimi-k2.7-code-highspeedkimi-k2

Quick Info

Powered by
Provider
Moonshot AI
Model key
kimi-k2.7-code-highspeed
Release date
Jun 12, 2026
Last updated
Jun 12, 2026
Knowledge cutoff
2025-01
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$1.90
Output token cost
$8.00

Limits

Output tokens
262,144 tokens
Context window
262,144 tokens

Transparent token rates

Compare Kimi K2.7 Code HighSpeed pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Kimi K2.7 Code HighSpeed

AIHubMix

CoverageBenchmark

The kimi-k2.7-code-highspeed variant is tuned for roughly 180 tokens/sec (up to 260 tok/s in short contexts), delivering roughly six times the throughput of the standard endpoint. It ships alongside kimi-k2.7-code, both released under a Modified MIT license that covers the model weights themselves, making this a genuin A notable design constraint is that thinking mode cannot be disabled on these models — every request runs the full chain-of-thought regardless of caller preference, and the API errors if users try to override temperature, top_p, or the penalty parameters away from their fixed defaults. Moonshot positions this as a deli

Ofox

Coverage

A third-party AI roundup dated June 16, 2026 explicitly names "Kimi K2.7 Code HighSpeed" and frames it not as a separate model but as a speed-optimized mode of the open-source Kimi K2.7 Code. It reports the same headline figures of approximately 180 tok/s on median-length coding tasks and peaks of up to 260 tok/s in sh The article attributes the speed gains to speculative decoding combined with optimized kernel scheduling tailored specifically for code-completion patterns, where token distributions are more constrained than in general chat. It reports early-adopter feedback of noticeable latency reductions in interactive coding sessi

Moonshot AI

Coverage

An EmpirioLabs API catalog entry dated June 16, 2026 documents Kimi K2.7 Code HighSpeed as the faster-serving tier of Moonshot's agentic coding model, with a 256K context window, always-on reasoning, multimodal inputs (text, image, and video), function calling, JSON Schema structured output, and a built-in web search t The same page lists live pay-as-you-go pricing of $1.90 per 1M prompt tokens and $8.00 per 1M generated tokens, with a 131,072 max output token limit and a release date of June 16, 2026, while noting that explicit cache controls, batching, and fine-tuning are not supported. Multi-step function calling requires the assi

Moonshot AI

Coverage

An AI news roundup dated June 13, 2026 reports that Moonshot AI open-sourced Kimi K2.7-Code as a coding-focused agentic model built on the K2.6 foundation, citing gains of 21.8% on Kimi Code Bench v2, 11.0% on Program Bench, and 31.5% on MLS Bench Lite, alongside roughly 30% fewer thinking tokens that translate into lo The same article places the Kimi release alongside unrelated announcements, including MiniMax M3 as a 428B-parameter sparse-attention open-weight model and Vercel's HarnessAgent for unified agent orchestration, and references industry commentary on the structural advantages closed-source APIs have in AI evaluations. Th

Moonshot AI

Official sourceDocumentation

Moonshot AI's platform documentation introduces Kimi K2.7 Code HighSpeed as a high-speed variant of the Kimi K2.7 Code coding model. The HighSpeed edition shares the same underlying model but outputs at approximately 180 tokens/s, reaching up to 260 tokens/s in short-context scenarios, offering a faster coding experience than the standard variant. The docs describe Kimi K2.7 Code HighSpeed as supporting a 256K context window, improved instruction compliance, and long-horizon coding versus the K2.6 predecessor, with Moonshot reporting roughly 30% less overthinking and 10% better agentic performance in external benchmarks. The API is OpenAI-compatible, no non-thinking mode is offered, and the HighSpeed tier is noted to be resource-constrained while Moonshot gradually expands capacity.

Ofox

Coverage

An encyclopedia entry explicitly titled "Kimi K2.7 Code Model High-Speed Edition" attributes the model to Moonshot AI ("Dark Side of the Moon") and dates the HighSpeed launch to June 15, 2026, three days after the standard K2.7 Code release on June 12, 2026. It reports the same ~180 tokens/s typical output speed and up The entry describes the initial access rollout as limited to Kimi Code Beta Program members, API developers, and Kimi Business users, with a planned gradual expansion to Kimi members at the Allegretto level and above. It also notes that thinking mode must be enabled to achieve optimal performance with the model series.

Moonshot AI

Coverage

Overchat's model directory page restates that Kimi K2.7 Code is a coding-focused open-weight flagship in Moonshot AI's Kimi K2 family, released on June 12, 2026 as the successor to K2.6 and shipped under a Modified MIT license with weights on Hugging Face at moonshotai/Kimi-K2.7-Code. It repeats the 1T-parameter MoE ar The same listing describes a separate kimi-k2.7-code-highspeed endpoint aimed at latency-sensitive agentic loops, citing the same ~180 tok/s typical and ~260 tok/s short-context throughput figures and a roughly sixfold speedup over the standard endpoint. It positions K2.7 Code as 5–12× cheaper per token than Claude Opu

GMI Cloud

Coverage

AgentsFlare reports that in June 2026 it added four new models: GLM-5.2, Kimi K2.7-Code standard and high-speed variants, and the GA version of Gemini 3 Pro Image. The update wave is concentrated around two capability tracks, coding agents and image generation. On the coding track, Moonshot released Kimi K2.7-Code and The AgentsFlare update states that Kimi K2.7-Code improves coding capability and reduces reasoning-token usage without raising its base price, positioning it competitively against GLM-5.2's approach of approaching Claude Opus 4.8-level engineering performance at lower cost. The post also flags retirement and deprecatio

Videos about Kimi K2.7 Code HighSpeed

More models around Kimi K2.7 Code HighSpeed