Currently listed through these providers:
Model details
MiniMax-M2.5-highspeed
MiniMax-M2.5-highspeed sits inside the MiniMax M-series lineup as a speed-tuned sibling to the standard M2.5 variant, sharing the same family identity while prioritizing quicker response throughput. The official API documentation lists it alongside the broader M-series, which spans earlier M2 and M2.1 releases as well as the more recent M2.7 and M3 generations, suggesting an iterative lineage focused on agentic reasoning, tool use, and long-context handling. Within that lineup, the highspeed label signals an emphasis on lower-latency generation rather than a different capability surface, making it a practical choice when interactive responsiveness matters more than squeezing the highest-quality output per token.
With a 204,800-token context window and a maximum output ceiling that comfortably accommodates extended generations, the model is well suited to long-document analysis, multi-turn agent loops, and tool-calling workflows that combine retrieval, code execution, and reasoning over large inputs. Its published capabilities center on reasoning, external tool calling, and temperature control, which together support classic agent patterns such as planning, function invocation, and reflective self-correction. Because it ships as open weights, teams that prefer self-hosting or fine-tuning can adapt it to domain-specific assistants without depending solely on hosted inference, which is a meaningful advantage for organizations with privacy or customization requirements.
Quick Info
Powered by- Provider
- MiniMax (minimax.io)
- Model key
- MiniMax-M2.5-highspeed
- Release date
- Feb 13, 2026
- Last updated
- Feb 13, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.60
- Output token cost
- $2.40
Limits
- Output tokens
- 131,072 tokens
- Context window
- 204,800 tokens