Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Fireworks AI logo

Model details

MiniMax-M3

MiniMax M3 is a multimodal foundation model that processes text, images, and video inputs while generating text output, designed specifically for sustained, multi-step development tasks rather than quick single-turn queries. Its defining architectural choice is MiniMax Sparse Attention (MSA), which replaces traditional full attention with a selective KV-block mechanism—a lightweight index branch scans incoming tokens and routes computation only to the most relevant key-value blocks. This approach achieves roughly one-twentieth the computational cost of the previous generation when operating at 1M tokens, delivering substantially faster prefill and decode without sacrificing output quality. The 1M-token context window makes it well-suited for reading large codebases, debugging complex projects, generating files across an entire repository, and handling intricate software development workflows that demand sustained attention across many files.

The model was trained as a natively multimodal system on interleaved data, meaning it learned to reason across text, images, and video from the ground up rather than bolting on vision capabilities afterward. Training incorporated an interactive user-simulator framework that tuned the model for multi-turn, production-like collaboration—essentially teaching it to maintain coherent reasoning across extended conversations. Reviewers have noted this generation marks a meaningful step forward from earlier M-series models, with one analyst suggesting the capability gap compared to leading proprietary models may have narrowed. The model is available with open weights, allowing developers to run it locally or through various hosted providers, and its combination of long-context capacity, multimodal understanding, and agentic reasoning positioning it for tooling ecosystems where sustained autonomous problem-solving matters more than isolated responses.

Fireworks AIaccounts/fireworks/models/minimax-m3minimax

Quick Info

Powered by
Provider
Fireworks AI
Model key
accounts/fireworks/models/minimax-m3
Release date
Jun 12, 2026
Last updated
Jun 12, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.30
Output token cost
$1.20

Limits

Output tokens
512,000 tokens
Context window
512,000 tokens

OpenCode

Model variants

Priority

Input
$0.45
per 1M tokens
Output
$1.80
per 1M tokens
Cache Read
$0.09
per 1M tokens

Transparent token rates

Compare MiniMax-M3 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about MiniMax-M3

Fireworks AI

Coverage

MiniMax officially released MiniMax-M3 on June 1, 2026, as described in the MiniMax Research blog. The model uses MSA (MiniMax Sparse Attention), a new sparse attention architecture introduced for M3, and supports an ultra-long context window of up to 1 million tokens. M3 is a natively multimodal model accepting image According to the official blog, M3 shows significant coding improvements over M2, approaching the level of leading overseas closed-source models in areas such as bugfix, frontend/backend development, and performance optimization. On agentic tasks, M3 performs strongly on common office workflows like search and Office-s

Videos about MiniMax-M3

More models around MiniMax-M3