Model details
Minimax/Minimax-M2.5 Highspeed
Minimax-M2.5 Highspeed is an accelerated variant of the standard M2.5, engineered to preserve the same reasoning depth and digital workspace capabilities while cutting response latency. Gateway and reseller descriptions frame it as inheriting the core intelligence of the base model, including its 80.2% score on SWE-Bench Verified, native Office document manipulation, and cross-software collaboration skills, then layering on engineering optimizations that yield an output speed of roughly 100 tokens per second. The design intent is to keep the high-quality planning and self-optimization behavior of the original model while turning it into a "high-velocity engine" suitable for latency-sensitive, high-frequency interactions.
In practical terms, the Highspeed profile is aimed at developers who need near real-time responsiveness without giving up serious code or document reasoning. The model is exposed through OpenAI-compatible serverless endpoints and a 204,800-token context window, making it well suited to interactive applications, large-scale automated pipelines, and complex document streams where both intelligence and throughput matter. Compared with a standard chat model, the trade-off is essentially transparent: the same core competencies as the parent M2.5, but tuned for scenarios where every millisecond of inference time compounds across many requests.
Quick Info
Powered by- Provider
- Qiniu
- Model key
- minimax/minimax-m2.5-highspeed
- Release date
- Feb 14, 2026
- Last updated
- Feb 14, 2026
- Input modalities
- Output modalities
- Capabilities
Limits
- Output tokens
- 128,000 tokens
- Context window
- 204,800 tokens