Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Vercel AI Gateway logo

Model details

MiniMax M2.5

MiniMax M2.5 is a large language model built for real-world productivity and coding tasks, designed to handle complex workflows that span full-stack development and office deliverables. The model is architected for agentic use cases, meaning it can decompose tasks, call tools, and reason through multi-step problems rather than just answering single prompts. Sources describe it as achieving state-of-the-art performance in coding, agentic tool use, search, and office work, positioning it as a practical workhorse rather than a research-only model.

The model was trained extensively using reinforcement learning operating across hundreds of thousands of machines, which MiniMax says enables efficient inference and optimized task decomposition at scale. This large-scale training environment gives the model exposure to complex, real-world conditions that smaller sandbox setups cannot replicate. A successor, M2.7, later extended these capabilities onto NVIDIA platforms for advanced agentic applications, suggesting the M2.5 lineage was intentionally designed to serve as a foundation for increasingly autonomous AI agents running in production environments.

Vercel AI Gatewayminimax/minimax-m2.5minimax

Quick Info

Powered by
Provider
Vercel AI Gateway
Model key
minimax/minimax-m2.5
Release date
Feb 12, 2026
Last updated
Feb 12, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.30
Output token cost
$1.20

Limits

Output tokens
131,000 tokens
Context window
204,800 tokens

Transparent token rates

Compare MiniMax M2.5 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about MiniMax M2.5

Merge Gateway

CoverageBenchmark

An InferenceX (SemiAnalysis) architecture profile covers MiniMax-M2.5 (and M2.7) as sparse MoE language models in MiniMax's M2 series, based on the technical report thesis that "mini activations can unleash maximum real-world intelligence." M2.5 is listed as released 2026-02-12 (matching MiniMax's announcement and Hugg The page documents specific architectural details: 62 MoE transformer layers, Grouped Query Attention with QK Norm, RoPE, Multi-Token Prediction (3 modules), Top-8/256 routing, RMSNorm, FP8 quantization, 197K context window, and d=3,072 token embeddings with a 200,064 vocabulary. It also cites M2.5's headline benchmark

Videos about MiniMax M2.5

More models around MiniMax M2.5