Sulat.com
AI models
Vercel AI Gateway logo

Model details

MiniMax M3

MiniMax M3 marks a deliberate course correction in the M-series lineage. Where the M2 generation stepped away from sparse attention over production concerns, M3 brings it back as the headline feature under the name MiniMax Sparse Attention. The mechanism works by using a lightweight index branch that scans incoming tokens and selects only the key-value blocks that actually warrant attention, running the expensive math on those specific blocks. Critically, this selection happens on real, uncompressed key-values, which the designers say sidesteps the typical long-context penalty that makes very large contexts impractical. Beyond the architectural shift, M3 is natively multimodal, accepting text, images, and video while producing text output, and it was clearly engineered with coding and agentic workflows as primary targets.

The model arrives with a benchmark claim that caught reviewers' attention: performance on SWE-bench competitive with GPT-5.5 and Opus-class models, a gap that earlier M-series releases had not closed. This positioning reflects a deliberate push toward practical coding tasks and multi-step agent behaviors rather than general-purpose chat. The open-weights approach means developers can run it locally or through API providers, and the sparse attention design is meant to make the cataloged API limit contexts genuinely usable for real work rather than a marketing headline. For teams building coding agents, long-context document reasoning, browser automation, or always-on assistant systems, M3 represents a version of the M-series that prioritizes production-ready agentic capability over incremental general intelligence gains.

Vercel AI Gatewayminimax/minimax-m3minimax

Quick Info

Powered by
Provider
Vercel AI Gateway
Model key
minimax/minimax-m3
Release date
Jun 1, 2026
Last updated
Jun 1, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.30
Output token cost
$1.20

Limits

Output tokens
512,000 tokens
Context window
512,000 tokens

Transparent token rates

Compare MiniMax M3 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about MiniMax M3

Vercel AI Gateway

Official sourceAnnouncement

MiniMax Research officially released MiniMax M3 on June 1, 2026, positioning it as the first and only open-weight model to combine frontier-level coding/agentic performance, native multimodality (image and video input), desktop-computer operation, and ultra-long context in a single model. The post claims significant co The release introduces MSA (MiniMax Sparse Attention), a new sparse attention architecture proposed by MiniMax's team that underpins the model's 1M-token context window by addressing the quadratic complexity of full attention. The blog details M3's availability through MiniMax Code, the Token Plan, and the MiniMax API,

Videos about MiniMax M3

More models around MiniMax M3