Mistral Large 3 marks Mistral AI's return to mixture-of-experts architecture after the Mixtral series, scaling up substantially as a sparse MoE with 41 billion active parameters drawn from 675 billion total parameters. It was trained from scratch on 3,000 NVIDIA H200 GPUs and released alongside three smaller dense Ministral variants (14B, 8B, and 3B) under the Apache 2.0 license, with the weights distributed in multiple compressed formats to support flexible deployment. After post-training, the model is positioned by Mistral as achieving frontier instruction-tuned quality among open-weight models on general prompts, while also delivering image understanding alongside its text I/O.
As Mistral's flagship open model, Mistral Large 3 is intended for developers and organizations that want frontier-class reasoning and multimodal input handling without licensing constraints, and its large parameter budget makes it well suited to complex instruction following, long-context workflows, and vision-plus-language tasks. The combination of open weights, permissive licensing, and frontier-tier post-training results makes it a practical choice for teams that need to self-host, fine-tune, or audit a capable general-purpose model, while smaller Ministral siblings offer lower-cost alternatives when raw scale is unnecessary.