Currently listed through these providers:
Model details
MiniMax M3 Fast
MiniMax M3 Fast sits inside the broader MiniMax M3 family and is positioned as an efficient-tier option alongside peers like GPT-5.6 Luna and DeepSeek V4 Flash in third-party Lite-tier comparisons. The family-level design uses a mixture-of-experts layout with roughly 428B total parameters and around 23B active per token, paired with MSA sparse attention to keep inference lean. That combination lets the fast variant serve as a lower-cost default for high-volume workloads where reasoning quality still matters but full flagship capacity is unnecessary.
Beyond cost efficiency, the family is built for native multimodal input: text, image, and video can enter the same call, while output stays text-only, which makes M3 Fast well suited to assistants that need to reason over mixed media such as documents, screenshots, or short clips. Open weights on Hugging Face add flexibility for teams that want to self-host or fine-tune, and independent comparison coverage notes roughly a 1M-token context window with the cataloged API limit guaranteed ceiling. In practice this profile fits routing pipelines where M3 Fast handles the bulk of requests cheaply, reserving heavier models for the hardest prompts.
Quick Info
Powered by- Provider
- Inco
- Model key
- minimax-m3:fast
- Release date
- Jun 1, 2026
- Last updated
- Jun 1, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.60
- Output token cost
- $2.40
Limits
- Output tokens
- 512,000 tokens
- Context window
- 1,048,576 tokens
Latest news about MiniMax M3 Fast
No articles yet. Fetch the latest news to show it here.