Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
ZenMux logo

Model details

MiniMax M2.5

MiniMax M2.5 is built as a 230B Mixture-of-Experts model that utilizes 10B active parameters to balance high-level performance with operational efficiency. Designed specifically for real-world productivity, the architecture is optimized for complex agentic tasks, including software engineering, advanced tool use, and search-based workflows. By focusing on efficient text generation, the model serves as a robust engine for office-style automation, providing a scalable solution for users who require consistent, high-quality reasoning across diverse and demanding professional environments.

The model benefits from advanced training methodologies, including reinforcement learning scaling and agent-native frameworks that enhance its ability to perform step-by-step verification. These techniques allow the model to function effectively as a professional employee, maintaining high performance while remaining cost-effective for continuous, long-running agentic operations. With support for modern inference backends and quantization, it is well-suited for self-hosted deployments and enterprise-grade applications that prioritize both accuracy and computational throughput.

ZenMuxminimax/minimax-m2.5

Quick Info

Powered by
Provider
ZenMux
Model key
minimax/minimax-m2.5
Release date
Feb 13, 2026
Last updated
Feb 13, 2026
Knowledge cutoff
2025-01-01
AI SDK package
@ai-sdk/anthropic
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.30
Output token cost
$1.20

Limits

Output tokens
131,072 tokens
Context window
204,800 tokens

Latest news about MiniMax M2.5

Merge Gateway

CoverageBenchmark

An InferenceX (SemiAnalysis) architecture profile covers MiniMax-M2.5 (and M2.7) as sparse MoE language models in MiniMax's M2 series, based on the technical report thesis that "mini activations can unleash maximum real-world intelligence." M2.5 is listed as released 2026-02-12 (matching MiniMax's announcement and Hugg The page documents specific architectural details: 62 MoE transformer layers, Grouped Query Attention with QK Norm, RoPE, Multi-Token Prediction (3 modules), Top-8/256 routing, RMSNorm, FP8 quantization, 197K context window, and d=3,072 token embeddings with a 200,064 vocabulary. It also cites M2.5's headline benchmark

Videos about MiniMax M2.5