Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
CoreWeave logo

Model details

MiniMax M3

MiniMax M3 is positioned as the flagship entry in MiniMax's large language model lineup, sitting alongside the M2.7 and M2.5 variants. The model is built around a native multimodal design rather than bolted-on adapters, with mixed-modality training applied from the very first training step. This early integration enables the model to develop deeper semantic fusion across text, image, and video inputs, supporting workflows that require reasoning over heterogeneous content in a single pass rather than relying on separate encoders stitched together after the fact.

Architecture-wise, M3 carries approximately 428 billion total parameters with roughly 23 billion activated per inference, suggesting a sparse mixture-of-experts style design that aims to balance capability against compute cost. To handle its million-token context window, the model introduces MiniMax Sparse Attention (MSA), a mechanism specifically engineered to keep long-context inference efficient. The open-weight release on Hugging Face, paired with availability through the MiniMax platform, makes M3 a practical choice for teams building long-context multimodal applications—such as video understanding, document analysis, or agentic systems that need to reason across large inputs—where both extended context capacity and vision-language grounding matter.

CoreWeaveMiniMaxAI/MiniMax-M3minimax-m3

Quick Info

Powered by
Provider
CoreWeave
Model key
MiniMaxAI/MiniMax-M3
Release date
Jun 12, 2026
Last updated
Jun 12, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.23
Output token cost
$0.96

Limits

Output tokens
262,144 tokens
Context window
262,144 tokens

Transparent token rates

Compare MiniMax M3 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about MiniMax M3

EmpirioLabs AI

Coverage

Verdent's engineering guide provides a builder-oriented read of MiniMax M3 (released June 1, 2026 by Shanghai-based MiniMax), framing it as a self-hostable, open-weight option for long-context agentic coding. The guide explicitly names M3 and cites the official MiniMax model card plus an arXiv technical report (arXiv:2 Concrete architectural details from the guide: M3 is a 428B-parameter Mixture-of-Experts model with roughly 22B parameters active per token, built on a Grouped-Query Attention backbone with MiniMax Sparse Attention (MSA) layered on top. The platform guarantees a minimum of 512K context (worth noting the official 1M fra

Merge Gateway

Coverage

Fireworks announced Day-0 support for MiniMax M3 on June 12, 2026, describing M3 as MiniMax's flagship frontier model with native multimodality, agentic capabilities, and a 1M-token context window in a single open-weight package. The post positions M3 against earlier 2026 open-weight releases, noting Kimi K2.5's Januar The launch post explains MiniMax Sparse Attention (MSA) as the long-context architecture underlying M3, and references the Artificial Analysis third-party intelligence index as showing M3 surpassing other open-source models in overall intelligence and exceeding several closed-source models including Opus 4.6. Fireworks

Merge Gateway

Coverage

Navneet Guglani's Medium analysis describes MiniMax M3 as released on June 1, 2026, claiming to be the first open-weight model to combine frontier-level coding performance, a 1-million-token context window, and native multimodal input in a single architecture. The API went live the same day with open weights and a tech The article traces MiniMax's M-series velocity — M1 in June 2025 with hybrid attention and 1M context, M2 in October 2025 with improved coding, M2.5 in February 2026 with stronger agentic tasks, and M2.7 in April 2026 with self-evolving agent capabilities — and argues M3 is positioned to cross into the closed-source fr

Merge Gateway

CoverageRelease Notes

MarkTechPost reports that MiniMax officially released MiniMax M3 on June 1, 2026, introducing MSA (MiniMax Sparse Attention) to provide a 1M-token context window along with native image and video input and desktop computer operation. The piece highlights M3 as the next model in the M-series line after M2.7 and notes th The article frames MSA as an architectural response to the quadratic compute scaling of full attention, framing M3 as a combined open-weight frontier model for coding, long context, and multimodality. It positions M3 against closed-source frontier models and emphasizes that open weights will enable enterprise customiza

Merge Gateway

CoverageBenchmark

VentureBeat reports that Chinese AI startup MiniMax released its MiniMax-M3 large language model on the evening of Sunday, June 1, 2026 (Eastern time), pairing frontier-tier coding and agentic performance with a 1-million-token context window and native multimodality. It quotes a launch-week API price of $0.30 per mill The piece frames M3 as combining previously separated frontier capabilities into a single open-weights system, with weights and a technical report promised within 10 days. It benchmarks M3 against MiMo-V2.5 Flash, DeepSeek V4 Flash/Pro, and Gemini 3.1 Flash-Lite on per-million-token cost, and discusses enterprise impli

CoreWeave

Coverage

MiniMax officially released MiniMax M3 on June 1, 2026, announcing it as the first open-weight model to combine frontier-level coding, a 1M token context window, and native multimodality (image and video input, plus desktop computer operation) in a single release. The official blog reports significant coding improvemen At the architectural core, MiniMax introduces MSA (MiniMax Sparse Attention), a new sparse attention design intended to escape the quadratic complexity of full attention and make context length a scalable dimension rather than a bottleneck. The blog frames MSA as a clean, extensible pre-filtering approach and contrasts

Merge Gateway

CoverageBenchmark

BenchLM's profile of MiniMax M3 (data as of September 22, 2026) lists it among 60 of 196 ranked public models, with a verified subset of 35 of 71. Quoted headline numbers include a capability score of 55.3/100 against a field median of 50.4, a 102 tok/s reported speed (field median 91 tok/s), a first-token latency of 2 Category-level results show Agentic at rank 45 of 105 (39.8 score), Coding at rank 63 of 135 (39.6 score), Multimodal at rank 36 of 50 (52.0 score), Knowledge at rank 62 of 160 (48.0 score), and Math measured at a single benchmark of 85.7 (not ranked). Reasoning, Multilingual, and Instruction Following carry weight but

Merge Gateway

Coverage

The Hugging Face model card for MiniMax-M3 describes a native multimodal model with 428B total parameters and 23B activated parameters, trained with mixed-modality data from the first step to fuse text, image, and video semantically. It introduces MiniMax Sparse Attention (MSA), which delivers 9× prefill and 15× decode Local deployment guidance on the card points developers at MiniMax Agent, the MiniMax API, and a downloadable weights repository, with files published in BF16 and F32 Safetensors formats. The page references the technical paper at arXiv:2606.13392 and reports 167,952 downloads in the last month prior to the scrape. Nat

Merge Gateway

CoverageBenchmark

OpenRouter's listing for MiniMax M3 describes it as a multimodal foundation model from MiniMax that supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding, and tool use. The page notes the model was released on May 31, 2026, and names Mini The listing details M3's training orientation as a native multimodal model on interleaved data, tuned for multi-turn production-like collaboration via an interactive user-simulator framework and oriented toward sustained multi-step tasks rather than single-turn execution. Modalities are summarized as text, image, and v

Videos about MiniMax M3

More models around MiniMax M3