Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Merge Gateway logo

Model details

MiniMax M2.5

MiniMax M2.5 is a Mixture-of-Experts model designed to deliver frontier-tier performance at a fraction of the typical compute cost. With 230 billion total parameters but only 10 billion activated during inference, the architecture keeps most of the model idle on any given pass, which is what makes the pricing viable. The model builds on the coding expertise of its predecessor M2.1 and extends into general office work, reaching fluency in generating and operating Word, Excel, and PowerPoint files, context switching between diverse software environments, and working across different agent and human teams. This full-stack agentic capability means M2.5 is not just a code model — it is positioned as a productivity workhorse that can handle tool calling, web search, and office workflows in a single session.

The model was trained using MiniMax's proprietary Forge reinforcement learning framework, which scales agent training across 200,000+ real-world environments including code repositories, web browsers, and office applications. The Forge framework uses a CISPO algorithm and achieves a 40x training speedup compared to earlier approaches. M2.5 scores 80.2% on SWE-Bench Verified — placing it within 0.6 percentage points of Claude Opus 4.6 — while also reaching 51.3% on Multi-SWE-Bench and 76.3% on BrowseComp. Both variants ship as open weights on Hugging Face under a modified MIT license, giving teams the freedom to self-host, fine-tune, or integrate the model directly into their pipelines. The Lightning variant doubles the standard speed to around 100 tokens per second, making it practical for real-time coding agents and high-throughput production workloads.

Merge Gatewayminimax/minimax-m2.5minimax

Quick Info

Powered by
Provider
Merge Gateway
Model key
minimax/minimax-m2.5
Release date
Feb 12, 2026
Last updated
Feb 12, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.30
Output token cost
$1.20

Limits

Output tokens
8,192 tokens
Context window
204,800 tokens

Transparent token rates

Compare MiniMax M2.5 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about MiniMax M2.5

Merge Gateway

CoverageBenchmark

An InferenceX (SemiAnalysis) architecture profile covers MiniMax-M2.5 (and M2.7) as sparse MoE language models in MiniMax's M2 series, based on the technical report thesis that "mini activations can unleash maximum real-world intelligence." M2.5 is listed as released 2026-02-12 (matching MiniMax's announcement and Hugg The page documents specific architectural details: 62 MoE transformer layers, Grouped Query Attention with QK Norm, RoPE, Multi-Token Prediction (3 modules), Top-8/256 routing, RMSNorm, FP8 quantization, 197K context window, and d=3,072 token embeddings with a 200,064 vocabulary. It also cites M2.5's headline benchmark

Videos about MiniMax M2.5

More models around MiniMax M2.5