NovitaAI
The release of MiniMax M2.7 adds enhancements to the popular MiniMax M2.5 model, built for agentic harnesses, and other complex use cases in fields such as…
Model details
MiniMax M2.5 is a large language model built from the ground up for agentic productivity, meaning it is designed to plan multi-step tasks, call tools, and move fluidly between different software environments rather than just produce single-turn answers. The model sits on a Mixture-of-Experts foundation with around 230 billion total parameters but only about 10 billion active per forward pass, a design that keeps inference cost low while preserving frontier-scale capability. It ships in two performance tiers, a standard variant running at roughly 50 tokens per second and a Lightning variant reaching about 100 tokens per second, and supports a very long context window with the underlying architecture extending to one million tokens. The intended sweet spot is real-world digital work: writing and refactoring code, searching the web, operating Word, Excel, and PowerPoint files, and coordinating across agent and human teams, with strong structured-output and tool-calling behavior built in.
The model extends the coding strengths of its predecessor M2.1 into broader office and general productivity territory, using a proprietary reinforcement learning framework called Forge that runs the model across more than 200,000 real-world environments spanning code repositories, browsers, and office applications. The Forge pipeline, which incorporates the CISPO algorithm and is reported to deliver a roughly 40x training speedup, optimizes how the model plans its actions and tokens rather than only chasing raw knowledge, producing measurable efficiency gains over earlier generations. This training focus is visible in the benchmark profile, where M2.5 reaches 80.2% on SWE-Bench Verified, 51.3% on Multi-SWE-Bench, 55.4% on SWE-Bench Pro, and 76.3% on BrowseComp, placing it within a small margin of leading proprietary systems at a fraction of the price. Released as open weights on Hugging Face under a modified MIT license, M2.5 is positioned as a flexible foundation for self-hosted coding agents, office automation, and multi-step workflows, while newer M2.7 updates build directly on the same agentic lineage.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
NovitaAI
The release of MiniMax M2.7 adds enhancements to the popular MiniMax M2.5 model, built for agentic harnesses, and other complex use cases in fields such as…
NovitaAI
Discover more about what's new at AWS with Minimax M2.5 and GLM 5 models now available on Amazon Bedrock
NovitaAI
Analyze MiniMax-M2.5 API latency, throughput, and cost efficiency benchmarks. Compare response speed, token performance, and pricing for scalable AI applications.
NovitaAI
MiniMax, an AI company based in Shanghai, China, has announced the MiniMax M2.5, a frontier model designed to dramatically improve real-world productivity. M2.5 uses reinforcement learning in complex real-world environments of hundreds of thousands of machines to achieve efficient inference and optimized task decomposi
Merge Gateway
An InferenceX (SemiAnalysis) architecture profile covers MiniMax-M2.5 (and M2.7) as sparse MoE language models in MiniMax's M2 series, based on the technical report thesis that "mini activations can unleash maximum real-world intelligence." M2.5 is listed as released 2026-02-12 (matching MiniMax's announcement and Hugg The page documents specific architectural details: 62 MoE transformer layers, Grouped Query Attention with QK Norm, RoPE, Multi-Token Prediction (3 modules), Top-8/256 routing, RMSNorm, FP8 quantization, 197K context window, and d=3,072 token embeddings with a 200,064 vocabulary. It also cites M2.5's headline benchmark