Model details
MiniMax-M2.5
MiniMax-M2.5 is a Mixture of Experts language model built to handle complex agentic workflows at scale. Its architecture pairs a large expert pool with selective activation, allowing the system to direct computational effort toward task-relevant parameters while keeping per-token inference lightweight. The design emphasis on real-world productivity means the model is particularly well-suited for software engineering tasks, structured tool use, search-heavy workflows, and multi-step office-style operations where reliable, coherent text generation matters most.
The model ships with FP8 quantization running on a SGLang inference backend, delivering efficient throughput across supported NVIDIA GPU platforms through a self-contained deployment container. An OpenAI-compatible API makes integration straightforward for teams that want to drop the model into existing pipelines without rethinking their client code. As an open-weight model, it invites teams to run, fine-tune, or extend it for domain-specific agentic applications where proprietary control or customization is a priority.
Quick Info
Powered by- Provider
- Nebius Token Factory
- Model key
- MiniMaxAI/MiniMax-M2.5
- Release date
- Jan 20, 2025
- Last updated
- May 7, 2026
- Knowledge cutoff
- 2025-01
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.30
- Output token cost
- $1.20
Limits
- Input tokens
- 190,000 tokens
- Output tokens
- 8,192 tokens
- Context window
- 196,608 tokens