Currently listed through these providers:
Model details
MiniMax M2.5 Highspeed
MiniMax M2.5 Highspeed is positioned within the MiniMax family as a streamlined, throughput-oriented variant of the broader M2.5 line, engineered for low-latency coding and agent-style workflows where response speed matters more than deep exploratory reasoning. Third-party directories describe it as an accelerated, high-throughput sibling of M2.5 that aims to preserve core intelligence and digital workspace capabilities while trimming overhead, making it a natural fit for interactive developer tools, code completion, tool-using assistants, and other latency-sensitive automation pipelines.
The model is catalogued as a text-to-text language model with an extended context window reaching roughly 205K tokens and a maximum output ceiling reported in the hundreds of thousands of tokens, paired with competitive per-token pricing that supports high-volume, agentic workloads. Capability tags from third-party listings highlight tool calling, reasoning, and temperature control, suggesting it is built to slot directly into orchestration frameworks that chain model calls, function calls, and structured reasoning steps. Practically, it suits teams that need a fast, affordable general-purpose LLM for coding copilots, retrieval-augmented assistants, and agent loops where round-trip time and cost per turn are primary constraints.
Quick Info
Powered by- Provider
- Merge Gateway
- Model key
- minimax/minimax-m2.5-highspeed
- Release date
- Feb 13, 2026
- Last updated
- Feb 13, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.60
- Output token cost
- $2.40
Limits
- Output tokens
- 8,192 tokens
- Context window
- 204,800 tokens
Transparent token rates
Compare MiniMax M2.5 Highspeed pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about MiniMax M2.5 Highspeed
No articles yet. Fetch the latest news to show it here.