Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
DevPass (LLM Gateway) logo

Model details

Kimi K2.7 Code Highspeed

Kimi K2.7 Code Highspeed is the speed-optimized tier of Moonshot's agentic coding model, purpose-built for development workflows that demand both deep reasoning and rapid response. It is described as the same underlying weights as Kimi K2.7 Code, but served with an accelerated inference profile that reaches roughly 180 tokens per second and up to about 260 tokens per second in short-context scenarios. That emphasis on throughput reflects an intent to support interactive coding sessions, long-running agent loops, and tooling-heavy tasks where latency directly shapes developer experience. The model carries a native multimodal architecture that takes in text, images, and video while emitting text, and operates with reasoning always on so each generation reflects deliberate chain-of-thought planning rather than reflexive pattern completion.

As the follow-up to K2.6 within the kimi-k2 family, Kimi K2.7 Code Highspeed is positioned around measurable gains in instruction compliance and long-horizon coding performance, with external benchmarks cited as showing a 30 percent average reduction in overthinking tendencies compared to the prior generation. It runs on a 256K context window, enabling it to hold substantial codebases, tool histories, and multi-step agent state in a single pass, and it supports structured output through JSON mode, function calling, and automatic context caching for cost-efficient repeated prefixes. The service uses fixed sampling settings, which means temperature and similar overrides are ignored in favor of deterministic behavior suited to reproducible agentic runs. Its practical sweet spot is extended, multi-file coding tasks that pair multimodal inputs with tool use, where faster token delivery translates into shorter wall-clock time without sacrificing the long-context reasoning that defines the K2.7 Code line.

DevPass (LLM Gateway)kimi-k2.7-code-highspeedkimi-k2

Quick Info

Powered by
Provider
DevPass (LLM Gateway)
Model key
kimi-k2.7-code-highspeed
Release date
Jun 12, 2026
Last updated
Jun 12, 2026
Knowledge cutoff
2025-01
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$1.90
Output token cost
$8.00

Limits

Output tokens
262,144 tokens
Context window
262,144 tokens

Transparent token rates

Compare Kimi K2.7 Code Highspeed pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Kimi K2.7 Code Highspeed

AIHubMix

CoverageBenchmark

The kimi-k2.7-code-highspeed variant is tuned for roughly 180 tokens/sec (up to 260 tok/s in short contexts), delivering roughly six times the throughput of the standard endpoint. It ships alongside kimi-k2.7-code, both released under a Modified MIT license that covers the model weights themselves, making this a genuin A notable design constraint is that thinking mode cannot be disabled on these models — every request runs the full chain-of-thought regardless of caller preference, and the API errors if users try to override temperature, top_p, or the penalty parameters away from their fixed defaults. Moonshot positions this as a deli

Ofox

Coverage

A third-party AI roundup dated June 16, 2026 explicitly names "Kimi K2.7 Code HighSpeed" and frames it not as a separate model but as a speed-optimized mode of the open-source Kimi K2.7 Code. It reports the same headline figures of approximately 180 tok/s on median-length coding tasks and peaks of up to 260 tok/s in sh The article attributes the speed gains to speculative decoding combined with optimized kernel scheduling tailored specifically for code-completion patterns, where token distributions are more constrained than in general chat. It reports early-adopter feedback of noticeable latency reductions in interactive coding sessi

Moonshot AI

Official sourceDocumentation

Moonshot AI's platform documentation introduces Kimi K2.7 Code HighSpeed as a high-speed variant of the Kimi K2.7 Code coding model. The HighSpeed edition shares the same underlying model but outputs at approximately 180 tokens/s, reaching up to 260 tokens/s in short-context scenarios, offering a faster coding experience than the standard variant. The docs describe Kimi K2.7 Code HighSpeed as supporting a 256K context window, improved instruction compliance, and long-horizon coding versus the K2.6 predecessor, with Moonshot reporting roughly 30% less overthinking and 10% better agentic performance in external benchmarks. The API is OpenAI-compatible, no non-thinking mode is offered, and the HighSpeed tier is noted to be resource-constrained while Moonshot gradually expands capacity.

Ofox

Coverage

An encyclopedia entry explicitly titled "Kimi K2.7 Code Model High-Speed Edition" attributes the model to Moonshot AI ("Dark Side of the Moon") and dates the HighSpeed launch to June 15, 2026, three days after the standard K2.7 Code release on June 12, 2026. It reports the same ~180 tokens/s typical output speed and up The entry describes the initial access rollout as limited to Kimi Code Beta Program members, API developers, and Kimi Business users, with a planned gradual expansion to Kimi members at the Allegretto level and above. It also notes that thinking mode must be enabled to achieve optimal performance with the model series.

Moonshot AI

Coverage

Overchat's model directory page restates that Kimi K2.7 Code is a coding-focused open-weight flagship in Moonshot AI's Kimi K2 family, released on June 12, 2026 as the successor to K2.6 and shipped under a Modified MIT license with weights on Hugging Face at moonshotai/Kimi-K2.7-Code. It repeats the 1T-parameter MoE ar The same listing describes a separate kimi-k2.7-code-highspeed endpoint aimed at latency-sensitive agentic loops, citing the same ~180 tok/s typical and ~260 tok/s short-context throughput figures and a roughly sixfold speedup over the standard endpoint. It positions K2.7 Code as 5–12× cheaper per token than Claude Opu

GMI Cloud

Coverage

AgentsFlare reports that in June 2026 it added four new models: GLM-5.2, Kimi K2.7-Code standard and high-speed variants, and the GA version of Gemini 3 Pro Image. The update wave is concentrated around two capability tracks, coding agents and image generation. On the coding track, Moonshot released Kimi K2.7-Code and The AgentsFlare update states that Kimi K2.7-Code improves coding capability and reduces reasoning-token usage without raising its base price, positioning it competitively against GLM-5.2's approach of approaching Claude Opus 4.8-level engineering performance at lower cost. The post also flags retirement and deprecatio

Videos about Kimi K2.7 Code Highspeed

More models around Kimi K2.7 Code Highspeed