Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Deep Infra logo

Model details

MiniMax M2.5

MiniMax M2.5 is built around a Mixture-of-Experts architecture that keeps the model efficient at inference. With 230 billion total parameters but only 10 billion active per pass, it delivers frontier-level capabilities without the usual compute overhead. The model ships in two variants—Standard and Lightning—catering to different throughput needs, and it is released as open weights on Hugging Face. Its design centers on code generation and refactoring, with polyglot code mastery and precision as core strengths.

The model was trained using Forge, MiniMax's proprietary reinforcement learning framework, deployed across more than 200,000 real-world environments. This training approach helped it achieve SWE-Bench Verified performance within 0.6% of Opus 4.6, at roughly one-twentieth the cost. As a lower-cost option within the MiniMax family, M2.5 fits coding agents, repository Q&A, research tasks, and fallback routes where budget matters. Teams can route by task value and context size, keeping M2.5 as a reliable default while escalating to M3 for more demanding agentic workflows.

Deep InfraMiniMaxAI/MiniMax-M2.5minimaxdeprecated

Quick Info

Powered by
Provider
Deep Infra
Model key
MiniMaxAI/MiniMax-M2.5
Release date
Feb 12, 2026
Last updated
Feb 12, 2026
Knowledge cutoff
2025-06
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.15
Output token cost
$1.15

Limits

Output tokens
131,072 tokens
Context window
196,608 tokens

Transparent token rates

Compare MiniMax M2.5 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about MiniMax M2.5

Merge Gateway

CoverageBenchmark

An InferenceX (SemiAnalysis) architecture profile covers MiniMax-M2.5 (and M2.7) as sparse MoE language models in MiniMax's M2 series, based on the technical report thesis that "mini activations can unleash maximum real-world intelligence." M2.5 is listed as released 2026-02-12 (matching MiniMax's announcement and Hugg The page documents specific architectural details: 62 MoE transformer layers, Grouped Query Attention with QK Norm, RoPE, Multi-Token Prediction (3 modules), Top-8/256 routing, RMSNorm, FP8 quantization, 197K context window, and d=3,072 token embeddings with a 200,064 vocabulary. It also cites M2.5's headline benchmark

Deep Infra

CoverageBenchmark

Third-party monitoring site LLM Benchmarks reports fresh performance data specifically for DeepInfra's MiniMax M2.5 endpoint, with the most recent run completed on August 4, 2026. Across 3 measured runs, the model averages 22.70 visible tokens per second (range 19.90–27.30) with a reported average time to first token o The page positions DeepInfra as one provider in a broader benchmark comparison, with throughput-distribution and time-series charts designed to normalise visible-token throughput (excluding reasoning tokens) so thinking-style models do not appear artificially slow. The data is useful as a recent developer-facing signal

Videos about MiniMax M2.5

More models around MiniMax M2.5