Sulat.com
AI models
Mistral logo

Model details

Mistral Large 3

Mistral Large 3 is a sparse mixture-of-experts language model that activates 41 billion parameters during inference while drawing on a total pool of 675 billion parameters. It marks Mistral's return to the MoE architecture after the Mixtral series, designed to deliver frontier-level performance across general instruction-following tasks. The model also incorporates vision capabilities, enabling it to process and reason about images alongside text. By structuring computation as a sparse mixture of specialized expert pathways, the architecture aims to deliver high capability at a lower active-parameter cost than a comparable dense model would require.

The model was trained from scratch using a cluster of 3000 NVIDIA H200 GPUs, representing a substantial investment in pretraining scale. After post-training, Mistral Large 3 achieves parity with the best instruction-tuned open-weight models on general prompts and demonstrates best-in-class per-token performance on image understanding tasks. Released under the permissive Apache 2.0 license and available in compressed formats, it is built for developers and organizations that want full control over deployment without licensing restrictions. The combination of open weights, vision support, and strong instruction-following makes it well suited for research, enterprise applications, and anyone building on open foundation models.

Mistralmistral-large-2512mistral-large

Quick Info

Powered by
Provider
Mistral
Model key
mistral-large-2512
Release date
Nov 1, 2024
Last updated
Dec 2, 2025
Knowledge cutoff
2024-11
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.50
Output token cost
$1.50

Limits

Output tokens
262,144 tokens
Context window
262,144 tokens

Latest news about Mistral Large 3

No articles yet. Fetch the latest news to show it here.

Videos about Mistral Large 3

More models around Mistral Large 3