Azure
Mistral Small 3.1 in May 2026: 128k context, vision, 80.6% MMLU, Apache 2.0. Plus where Small 3.2, Medium 3, and Mistral Large 2 fit the lineup.
Model details
Mistral Small 3.1 is a compact, multimodal large language model positioned as a versatile generalist that pairs text understanding with vision capabilities in a single 24-billion-parameter package. Building directly on the earlier Mistral Small 3, it introduces state-of-the-art image understanding while extending the context window far beyond typical small-model limits, aiming to handle long documents, technical reasoning, and conversational tasks without sacrificing raw text quality. The model is engineered for low-latency interactive use, with reported inference speeds around 150 tokens per second and an instruction-tuned design that targets chat, programming assistance, mathematical problem solving, and document comprehension. Its multimodal design lets it process both text and visual inputs together, while a multilingual training base spanning dozens of languages broadens its usefulness for global applications.
The release represents an evolution of the Mistral Small family rather than a wholesale retrain, refining the prior version with improved text performance, vision understanding, and longer-context behavior, and it is published as an instruction-finetuned variant of a public pretrained base. The accompanying material highlights strong results on reasoning benchmarks such as GPQA Diamond, knowledge suites including MMLU and MMLU-Pro, reading comprehension like TriviaQA, and quantitative and coding tasks such as MATH and HumanEval, where it is positioned as outperforming comparable small proprietary systems including Gemma 3 and GPT-4o Mini. Native function calling and structured JSON output give it agent-ready behavior suited for fast-response assistants and low-latency tool use, and its compact parameter count makes it practical for local inference on a single high-end consumer GPU or a quantized laptop setup. The Apache 2.0 licensing of the underlying weights, combined with strong baseline quality, makes it a flexible foundation for domain-specific fine-tuning, private on-device deployments, and enterprise workflows that need sensitive data to remain local.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Azure
Mistral Small 3.1 in May 2026: 128k context, vision, 80.6% MMLU, Apache 2.0. Plus where Small 3.2, Medium 3, and Mistral Large 2 fit the lineup.