Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
ZenMux logo

Model details

MiniMax-M3

MiniMax-M3 is positioned as a native multimodal flagship large language model, distinguished by being trained with mixed text, image, and video data from the initial pre-training stage rather than having modalities bolted on later. This early-stage multimodal training approach is paired with a sparse attention architecture, which the development team has documented in a dedicated technical paper released alongside the open-weight launch. Together, these design choices aim to give the model coherent cross-modal understanding while keeping inference costs manageable through selective activation of its parameters.

The model's standout practical strengths are its coding and agentic capabilities, where it has delivered industry-leading results on high-difficulty evaluations and rose quickly to the top of open-source rankings on a major global intelligence index shortly after release. Response speed was also a focal improvement area, with the team raising output throughput from an initial 30 TPS to 80 TPS and signaling further optimizations ahead. These traits make MiniMax-M3 a strong fit for developers building complex coding assistants, autonomous agents, and other reasoning-heavy applications that benefit from open weights and long-context multimodal input.

ZenMuxminimax/minimax-m3minimax

Quick Info

Powered by
Provider
ZenMux
Model key
minimax/minimax-m3
Release date
Jun 1, 2026
Last updated
Jun 1, 2026
AI SDK package
@ai-sdk/anthropic
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.60
Output token cost
$2.40

Limits

Output tokens
512,000 tokens
Context window
1,048,576 tokens

Transparent token rates

Compare MiniMax-M3 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about MiniMax-M3

Ollama Cloud

CoverageRelease Notes

A release tracker confirms MiniMax M3 was released on June 1, 2026, as the newest entry in the MiniMax model lineup. The tracker lists M3 with a 512K context window, input pricing of $0.30 per million tokens, and output pricing of $1.20 per million tokens, with an intelligence score of 45. This date corroborates the of The tracker lists six prior MiniMax releases dating back to October 2025, showing the progression of the M-series from M1 through M2 and its variants to M3. The 512K context figure shown here differs from the 1M context reported by the vendor and other third-party analyses, likely reflecting a served variant or aggrega

Ollama Cloud

CoverageBenchmark

A third-party technical breakdown confirms MiniMax M3 as a 428-billion-parameter mixture-of-experts model with approximately 23 billion active parameters per token, released June 1, 2026. The analysis details that the attention stack combines grouped-query attention with MiniMax Sparse Attention (MSA), which reportedly The breakdown reports that model weights went live on Hugging Face by June 7, 2026, and the technical report was published on arXiv on June 11, 2026, under a custom MiniMax-community license. It also notes an interleaved-thinking tool-calling behavior as a practical integration consideration. Official pricing is listed

Volcengine Ark Coding Plan

CoverageBenchmark

An independent technical deep-dive (July 5, 2026) details the architecture behind MiniMax-M3, describing it as a Mixture-of-Experts model with roughly 428B total parameters and about 23B active per token, paired with MiniMax Sparse Attention (MSA). MSA is explained as a two-branch design: a lightweight Index Branch tha Beyond architecture, the piece contrasts benchmark reality against the marketing claims and walks through production integration via an OpenAI-compatible unified gateway. For developers evaluating M3, the core takeaway is that long-context efficiency is real and hardware-dependent (best on H800-class GPUs), while the g

302.AI

Coverage

A Verdent developer guide dated June 15, 2026, cross-referenced against MiniMax's official model card and the arXiv technical report (arXiv:2606.13392), provides builder-oriented architectural detail on MiniMax M3. The page describes M3 as a 428B-parameter Mixture-of-Experts model with approximately 22B parameters acti The guide flags a non-standard license as a caveat often skipped in other coverage, and emphasizes that M3's headline metrics — 1M context, multimodality, and open weights — need to be translated into engineering decisions around long-repo reading, tool-calling loops, and multi-step agentic tasks. It positions MiniMax'

OpenRouter

CoverageBenchmark

An independent benchmark comparison (ofox.ai, June 15, 2026) reports MiniMax-M3 scoring 59.0% on SWE-Bench Pro versus GPT-5.5's 58.6%, with M3 priced at $0.60 input / $2.40 output per million tokens compared to GPT-5.5's $5 / $30 — roughly 8–12× cheaper per SWE-Bench Pro point. The post frames M3 as the first open-weig The comparison notes GPT-5.5 retains advantages on Terminal-Bench (82.7% vs M3's 66.0%) and Codex CLI integration, while M3 is favored for long-context refactors (1M context with MSA, ~15× faster decode than M2.5), native multimodal code review, and air-gapped on-prem deployment via open weights on Hugging Face.

OpenRouter

Coverage

Fireworks AI published a launch post (June 12, 2026) announcing Day-0 support for MiniMax-M3, supporting up to 500K tokens of context at launch and partnering with MiniMax to bring the full 1M-token window shortly after. The post frames M3 as MiniMax's flagship frontier model — a 500K-token context open-weight release The post cites Artificial Analysis' third-party intelligence index, stating M3 surpasses all other open-source models in overall intelligence and exceeds several closed-source models including Opus 4.6. (Note: this candidate is largely a serving-provider announcement, so it is down-weighted but still useful to corrobor

DevPass (LLM Gateway)

CoverageAnalysis

Artificial Analysis published an independent third-party evaluation on June 8, 2026, scoring MiniMax-M3 at 55 on the Artificial Analysis Intelligence Index, placing it ahead of open-weights peers Kimi K2.6 (54) and MiMo-V2.5-Pro (54). M3 is MiniMax's first multimodal M-series model, adding image/video input and a 1M-to Specific benchmark improvements over M2.7 include HLE +9 (28%→37%), GPQA Diamond +6 (87%→93%), AA-LCR +5 (69%→74%), IFBench +7 (76%→83%), and CritPt +3 (1%→4%), with a small SciCode regression (47%→45%). M3 scores 1670 on GDPval-AA (behind Claude Opus 4.8 max at 1890 and GPT-5.5 xhigh at 1769, level with Claude Sonnet

Inco

Coverage

Navneet Guglani's Medium analysis provides third-party technical context for MiniMax-M3, confirming its June 1, 2026 launch and its status as the first open-weight model to combine frontier-level coding, a 1-million-token context window, and native multimodal input in one architecture. The piece traces MiniMax's rapid Beyond launch details, the article notes that the M3 API went live on the same day as the release, with open weights and a technical report scheduled within 10 days. It positions M3 against closed-source frontier models such as Claude Opus, GPT-5.5, and Gemini 3.1 Pro, arguing that while open models have previously exc

Volcengine Ark Coding Plan

Coverage

An independent Stackademic hands-on test (June 3, 2026) puts MiniMax M3 through a full day of real Python and agentic workloads against Claude Opus 4.8 and GPT-5.5, using identical prompts, codebases, and workflows. The reviewer reports M3 scoring roughly 59% on SWE-Bench Pro (close to Claude Opus 4.8), strong performa The write-up describes M3 as an open-weights model built specifically for coding, long-context work, and AI agents, with its sparse attention mechanism enabling the 1M-token context window by focusing compute on the most relevant prompt regions rather than every token. While useful as a third-party developer signal on

MiniMax (minimax.cn)

Coverage

MiniMax officially released MiniMax M3 on June 1, 2026, describing it as a next-generation general-purpose model that simultaneously reaches frontier-level performance on coding and agentic tasks, supports ultra-long context windows of up to 1M tokens, and delivers native multimodality with image and video input plus d The central architectural innovation enabling M3's 1M context scaling is MSA (MiniMax Sparse Attention), a new attention mechanism proposed by MiniMax that addresses the quadratic complexity problem inherent in full attention. The blog describes MSA as a clean, easily extensible architecture that makes context "truly a

MiniMax (minimax.io)

CoverageBenchmark

The llm-stats.com model page tracks MiniMax-M3 across an aggregated leaderboard, ranking it #50 overall on the composite LLM Stats Score, with category standings including Coding 39 of 266, Vision 43 of 208, Reasoning 48 of 362, Math 107 of 327, and Finance 65 of 227, plus a tier-B overall standing. It also reports a b Quality Tracker data on the page lists 187 votes over the prior seven days and a +0.87σ stable quality signal, with sub-category moves including Chat +1.31σ (119 votes) and Websites +1.33σ (38 votes), while conversation-depth scores dip from 15.7 at turn 1 to 13.2 at turns 31+, a decline of −1.9σ from turn 1. These fig

Vancine

CoverageBenchmark

BenchLM's tracker page for MiniMax M3, dated as of September 15, 2026, records the model as released on June 1, 2026 with open-weight licensing, a 1M-token context window, and a non-reasoning profile. It reports a composite score of 61.5/100, ranking 54 of 232 tracked models, with the strongest eligible category being The BenchLM entry surfaces 27 published benchmark rows out of the tracked set, leaving some benchmark slots empty, and frames M3 as a model whose competitive strength is concentrated in multimodal-grounded workflows (screenshots, documents, charts) while its agentic and coding placements sit mid-pack. The page is an in

Videos about MiniMax-M3

More models around MiniMax-M3