Currently listed through these providers:
Model details
MiniMax M2.5 High Speed
MiniMax M2.5 High Speed is a high-throughput variant engineered for applications that need both speed and intelligence. The model delivers three times faster inference compared to its base counterpart while maintaining the core planning and software engineering capabilities of the M2.5 family. With a 205K token context window and support for up to 131K output tokens, it handles extended reasoning tasks and complex chain-of-thought problems without sacrificing depth. The combination of extreme efficiency with full capability preservation makes it suitable for production environments where latency matters.
The model quickly gained traction after launch, being integrated into over 50 platforms, suggesting strong practical demand for its balanced approach. Its native tool calling and parallel function calling capabilities enable it to interact with external systems and APIs seamlessly. The extensive suite of features including adaptive reasoning, prompt caching, code execution, and structured outputs positions it well for modern AI application development. The focus on maintaining full intelligence while optimizing throughput reflects a maturation in how AI models are being designed for real-world deployment.
Quick Info
Powered by- Provider
- Vercel AI Gateway
- Model key
- minimax/minimax-m2.5-highspeed
- Release date
- Feb 13, 2026
- Last updated
- Feb 13, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.60
- Output token cost
- $2.40
Limits
- Output tokens
- 131,000 tokens
- Context window
- 204,800 tokens
Transparent token rates
Compare MiniMax M2.5 High Speed pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about MiniMax M2.5 High Speed
No articles yet. Fetch the latest news to show it here.