Currently listed through these providers:
Model details
MiniMax-M2.5-highspeed
MiniMax-M2.5-highspeed is a performance-optimized variant engineered to deliver the same core capabilities as the standard M2.5 model while dramatically accelerating inference throughput. The design intent centers on efficiency without compromise: it preserves polyglot code mastery, precision code refactoring, and the full reasoning toolkit that developers rely on, but accelerates output to approximately 100 tokens per second—roughly three times faster than the base M2.5 variant. This makes it particularly well-suited for interactive coding environments, real-time agentic workflows, and applications where low latency directly impacts user experience or throughput economics.
The M2.5 family saw rapid adoption after launch, quickly integrating into over 50 platforms, which underscored demand for high-speed inference in production coding scenarios. The highspeed variant emerged as a direct response to that market feedback, targeting developers who need the intelligence of a frontier code model but cannot tolerate the latency of standard inference. Being open weights and available on HuggingFace, it invites community experimentation and self-hosting, extending MiniMax's reach beyond API-only offerings. For teams building code agents, automated refactoring pipelines, or any software tool that demands both deep reasoning and snappy response times, this model occupies a practical niche: a full-strength coding assistant that keeps pace with human-level iteration speed.
Quick Info
Powered by- Model key
- MiniMax-M2.5-highspeed
- Release date
- Feb 13, 2026
- Last updated
- Feb 13, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
A provider subscription or plan supersedes token-based pricing for this model.
Limits
- Output tokens
- 131,072 tokens
- Context window
- 204,800 tokens
Latest news about MiniMax-M2.5-highspeed
No articles yet. Fetch the latest news to show it here.