Currently listed through these providers:
Model details
Mistral Large 3
Mistral Large 3 is a sparse mixture-of-experts language model that activates 41 billion parameters during inference while drawing on a total pool of 675 billion parameters. It marks Mistral's return to the MoE architecture after the Mixtral series, designed to deliver frontier-level performance across general instruction-following tasks. The model also incorporates vision capabilities, enabling it to process and reason about images alongside text. By structuring computation as a sparse mixture of specialized expert pathways, the architecture aims to deliver high capability at a lower active-parameter cost than a comparable dense model would require.
The model was trained from scratch using a cluster of 3000 NVIDIA H200 GPUs, representing a substantial investment in pretraining scale. After post-training, Mistral Large 3 achieves parity with the best instruction-tuned open-weight models on general prompts and demonstrates best-in-class per-token performance on image understanding tasks. Released under the permissive Apache 2.0 license and available in compressed formats, it is built for developers and organizations that want full control over deployment without licensing restrictions. The combination of open weights, vision support, and strong instruction-following makes it well suited for research, enterprise applications, and anyone building on open foundation models.
Quick Info
Powered by- Provider
- Mistral
- Model key
- mistral-large-2512
- Release date
- Nov 1, 2024
- Last updated
- Dec 2, 2025
- Knowledge cutoff
- 2024-11
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.50
- Output token cost
- $1.50
Limits
- Output tokens
- 262,144 tokens
- Context window
- 262,144 tokens
Latest news about Mistral Large 3
No articles yet. Fetch the latest news to show it here.
Videos about Mistral Large 3
More models around Mistral Large 3
This exact model name is also listed by 11 other providers.