Model details
MiniMax-M3
MiniMax M3 is a multimodal foundation model that processes text, images, and video inputs while generating text output, designed specifically for sustained, multi-step development tasks rather than quick single-turn queries. Its defining architectural choice is MiniMax Sparse Attention (MSA), which replaces traditional full attention with a selective KV-block mechanism—a lightweight index branch scans incoming tokens and routes computation only to the most relevant key-value blocks. This approach achieves roughly one-twentieth the computational cost of the previous generation when operating at 1M tokens, delivering substantially faster prefill and decode without sacrificing output quality. The 1M-token context window makes it well-suited for reading large codebases, debugging complex projects, generating files across an entire repository, and handling intricate software development workflows that demand sustained attention across many files.
The model was trained as a natively multimodal system on interleaved data, meaning it learned to reason across text, images, and video from the ground up rather than bolting on vision capabilities afterward. Training incorporated an interactive user-simulator framework that tuned the model for multi-turn, production-like collaboration—essentially teaching it to maintain coherent reasoning across extended conversations. Reviewers have noted this generation marks a meaningful step forward from earlier M-series models, with one analyst suggesting the capability gap compared to leading proprietary models may have narrowed. The model is available with open weights, allowing developers to run it locally or through various hosted providers, and its combination of long-context capacity, multimodal understanding, and agentic reasoning positioning it for tooling ecosystems where sustained autonomous problem-solving matters more than isolated responses.
Quick Info
Powered by- Provider
- Fireworks AI
- Model key
- accounts/fireworks/models/minimax-m3
- Release date
- Jun 12, 2026
- Last updated
- Jun 12, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.30
- Output token cost
- $1.20
Limits
- Output tokens
- 512,000 tokens
- Context window
- 512,000 tokens
OpenCode
Model variants
Transparent token rates
Compare MiniMax-M3 pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about MiniMax-M3
Fireworks AI
MiniMax officially released MiniMax-M3 on June 1, 2026, as described in the MiniMax Research blog. The model uses MSA (MiniMax Sparse Attention), a new sparse attention architecture introduced for M3, and supports an ultra-long context window of up to 1 million tokens. M3 is a natively multimodal model accepting image According to the official blog, M3 shows significant coding improvements over M2, approaching the level of leading overseas closed-source models in areas such as bugfix, frontend/backend development, and performance optimization. On agentic tasks, M3 performs strongly on common office workflows like search and Office-s
Videos about MiniMax-M3
More models around MiniMax-M3
This exact model name is also listed by 42 other providers.