Currently listed through these providers:
Model details
MiniMax-M3
MiniMax-M3 belongs to the minimax family of large language models and is positioned as a native multimodal system, with the public product page describing it as pre-trained on over 100 trillion tokens of interleaved multimedia data so that text, images, and video share a unified representation space. The same developer-forum discussion frames the model as a target for NVFP4 quantization on Quad DGX Spark hardware, suggesting that the weights are amenable to low-precision inference on consumer and prosumer NVIDIA silicon. As an open-weight release, it gives researchers and builders direct access to the parameters rather than only a hosted endpoint, which makes it attractive for self-hosted experimentation, fine-tuning, and reproducible benchmarks. In practical terms, MiniMax-M3 is suited to workloads that blend modalities and benefit from very large context windows, including document and chart understanding, video question answering, and agentic pipelines that need to reason across long transcripts or mixed attachments. The presence of an active community thread focused on deploying it via NVFP4 on DGX Spark indicates real interest in running it efficiently on edge or workstation-class GPUs, while its placement alongside M2.7 and M2.5 in the provider's LLM lineup signals that it is meant to be a forward step in the series. Teams looking for a flexible open-weight multimodal foundation model with quantization-friendly deployment characteristics will find it a reasonable fit, especially when long-context reasoning and tool-driven workflows are central requirements.
MiniMax-M3 belongs to the minimax family of large language models and is positioned as a native multimodal system that learns directly from interleaved text, image, and video data, a design that lets a single checkpoint reason across formats without bolted-on encoders. The public model page describes pre-training on more than 100 trillion tokens, a scale that supports broad world knowledge and cross-modal grounding, while community activity on NVIDIA's developer forums demonstrates concrete interest in deploying the weights with NVFP4 quantization on Quad DGX Spark hardware for efficient local inference. Sibling entries M2.7 and M2.5 in the same product lineup confirm an iterative development trajectory, with M3 representing the most capable tier in the series. Practically, the model is a strong match for long-context agentic tasks such as multimodal document analysis, video understanding, and tool-augmented assistants that must hold large transcripts or attachment payloads in memory. Open-weight availability lowers the barrier for self-hosting, fine-tuning, and quantization research, while the confirmed DGX Spark deployment interest suggests predictable behavior on mainstream NVIDIA GPUs. Teams that need a unified text-plus-vision-and-video foundation model with an active deployment ecosystem should find MiniMax-M3 a practical and forward-looking choice.
Quick Info
Powered by- Provider
- Abacus
- Model key
- MiniMaxAI/MiniMax-M3
- Release date
- Jun 1, 2026
- Last updated
- Jun 1, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.30
- Output token cost
- $1.20
Limits
- Output tokens
- 131,072 tokens
- Context window
- 1,000,000 tokens
Transparent token rates
Compare minimax pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about MiniMax-M3
Videos about MiniMax-M3
More models around MiniMax-M3
This exact model name is also listed by 39 other providers.