AIHubMix
The kimi-k2.7-code-highspeed variant is tuned for roughly 180 tokens/sec (up to 260 tok/s in short contexts), delivering roughly six times the throughput of the standard endpoint. It ships alongside kimi-k2.7-code, both released under a Modified MIT license that covers the model weights themselves, making this a genuin A notable design constraint is that thinking mode cannot be disabled on these models — every request runs the full chain-of-thought regardless of caller preference, and the API errors if users try to override temperature, top_p, or the penalty parameters away from their fixed defaults. Moonshot positions this as a deli