Mistral Large 4 represents the company's newest flagship general-purpose model, building on the Mistral Large lineage with a shift away from dense transformer design toward a sparse Mixture-of-Experts architecture. This configuration reportedly carries roughly 1.05 trillion total parameters while activating only about 49 billion during inference, paired with a 1.6 billion parameter vision encoder that supports its multimodal capabilities. The expanded context length, reaching one million tokens compared to the prior generation's 256K, positions the model for tasks involving long documents, extended codebases, and large-scale retrieval workflows where maintaining coherence over very long inputs matters.
The model is offered under a public preview status, with Mistral planning staged public weight releases following the initial API availability. Early launch benchmarks suggest competitive performance across coding, agentic, legal, and vision-grounded tasks, including reported scores on DeepSWE, Finch/FinWorkBench, Harvey Legal Agent, DIOR-RSVG, and CyBench, though these figures remain preliminary. Practically, Mistral Large 4 fits users who need a single frontier-tier system that can reason over very long contexts while accepting both text and image inputs, especially in enterprise search, legal analysis, and complex agent pipelines where deep context handling and multimodal grounding are equally important.