MiniMax Coding Plan (minimax.io)
After the release of the MiniMax M2.5 model, it was quickly integrated into over 50 platforms, and the M2.5-highspeed model was launched, with an inference spee
Model details
MiniMax-M2.5-highspeed is the accelerated sibling of the MiniMax-M2.5 coding model, released in February 2026 as part of the same M2 generation. MiniMax's own models documentation positions M2.5 and its highspeed variant as code-focused releases alongside the newer M2.7 line and the frontier M3, with M2.5 now appearing in the Legacy Models section of the index. Coverage from AIBase describes the highspeed variant as offering roughly three times the inference speed of the base M2.5 while keeping the same underlying performance, a trade-off that targets production coding assistants and agent pipelines where latency matters more than squeezing out the last few points of quality. The intended use is therefore practical rather than research-oriented: it is meant for developers and integrators who want a familiar M2.5 experience but cannot tolerate the slower response times of the standard variant.
On capability, the gateway listing from Merge shows that the model exposes text input and text output, with a tool-use output modality and explicit support for tool calling, tool choice, and streaming, plus structured output, which is a useful combination for agentic coding workflows. The same listing reports a 205K token context window and an 8K maximum output per response, with zero data retention enabled on the host. Pricing through the gateway is listed at $0.60 per million input tokens and $2.40 per million output tokens in USD, giving integrators a concrete sense of the cost profile for high-throughput agent traffic. In practice, this combination of a large context window, explicit tool-use support, and substantially faster inference makes the model a strong fit for interactive code completion, repository-level refactoring tasks, and multi-step agent loops where round-trip time is the main bottleneck.
A provider subscription or plan supersedes token-based pricing for this model.
MiniMax Coding Plan (minimax.io)
After the release of the MiniMax M2.5 model, it was quickly integrated into over 50 platforms, and the M2.5-highspeed model was launched, with an inference spee