Sulat.com
AI models
Wafer logo

Model details

MiniMax-M3

MiniMax-M3 sits at the top of the MiniMax text/LLM lineup, presented in the official API documentation as a current-generation model and announced through the MiniMax research blog under the headline "Frontier Coding, 1M that quick-info value, Native Multimodality — All in One Model." That framing positions the release as a step beyond earlier MiniMax-M2 entries, which were tuned around code generation and refactoring with a smaller activated footprint, into a model that combines native multimodality with a very large that quick-info value window for repository- and codebase-scale reasoning. The documentation row explicitly tags MiniMax-M3 as a "Frontier multimodal coding model with 1M that quick-info value window" and lists multimodal input, the 1M that quick-info value window, and frontier coding among its headline features, while the surrounding LLM navigation on the MiniMax site surfaces a dedicated model card page for further technical detail.

In practical terms, MiniMax-M3 is aimed at developers and engineering teams that need a single model to read large codebases, work with mixed text and visual inputs such as diagrams or screenshots, and drive tool-using coding agents end-to-end. The model is reachable both through the MiniMax platform and through Wafer's serverless inference layer, which advertises hosting of MiniMax-M3 alongside other open frontier models and provides setup paths for popular coding agents including Claude Code, Codex, Cline, Roo Code, Kilo Code, and OpenHands. That combination of long that quick-info value, native multimodality, and ready-made agent integrations makes MiniMax-M3 a natural fit for agentic software engineering workflows where sustained reasoning across large repositories matters more than single-prompt chat quality.

WaferMiniMax-M3minimax

Quick Info

Powered by
Provider
Wafer
Model key
MiniMax-M3
Release date
Jun 1, 2026
Last updated
Jun 1, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.33
Output token cost
$1.32

Limits

Output tokens
512,000 tokens
Context window
1,048,576 tokens

Transparent token rates

Compare MiniMax-M3 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about MiniMax-M3

Wafer

CoverageBenchmark

Independent technical breakdown confirms MiniMax M3 is a 428B-parameter mixture-of-experts model with roughly 23B active parameters per token, paired with a 1M-token context window and native image and video input trained from step zero. The attention stack uses grouped-query attention combined with MiniMax Sparse Atte M3 scores 80.5% on SWE-bench Verified at official pricing of $0.30 per million input tokens, making it among the cheapest models above 80% on that benchmark and the only one in its class that reads image and video natively. It matches Claude Sonnet 4.6 on real-world agentic benchmarks. For UI automation and multimodal

Wafer

Coverage

Independent analyst commentary frames MiniMax M3's June 1, 2026 release as a landmark for open-weight models because it is the first to combine frontier-level coding performance, a 1M-token context window, and native multimodal input in a single architecture. The API went live the same day as the announcement, with ope The analysis flags an important credibility caveat: model weights were not available at launch, and MiniMax's benchmarks are vendor-reported rather than independently verified. While MiniMax is described as a Shanghai-based AI company founded in late 2021 by former SenseTime executives, the article argues that the real

Wafer

CoverageBenchmark

The BenchLM benchmark aggregator tracks MiniMax M3 with a composite capability score of 63.9 out of 100, ranking it 49th of 232 models as of September 4, 2026. The model is classified as open-weight, non-reasoning, with a 1M-token context window. API pricing is $0.30 input and $1.20 output per million tokens, with cach Across 27 tracked benchmark rows, M3's strongest eligible category is Multimodal and Grounded, ranking 34th, with particular strength on screenshots, documents, charts, and grounded multimodal workflows. Category percentiles include Coding at the 63rd percentile and Knowledge at the 65th, while Agentic and Multimodal s

Wafer

Coverage

MiniMax officially released MiniMax M3 on June 1, 2026, as their flagship open-weight frontier model. The model uses a new attention architecture called MiniMax Sparse Attention (MSA), which the company describes as a clean, extensible sparse attention design that makes context another axis that can be scaled. It suppo On coding and agentic tasks, M3 shows significant improvements over the prior MiniMax M2, approaching leading closed-source models on bugfix, frontend/backend development, and performance optimization. It performs strongly on common office workflows such as search and Office-suite tasks and has become initially usable

Videos about MiniMax-M3

More models around MiniMax-M3