Mistral Small is a compact, low-latency language model from Mistral AI designed for efficient text generation and understanding. Positioned as a practical, accessible option within the Mistral family, it balances capability with computational efficiency. Early catalog records indicate it was built with a 32,000-token context window and released under commercial-use terms, making it suitable for developers and businesses seeking a capable but resource-friendly model for chat, analysis, and general-purpose language tasks.
The Mistral Small family has evolved significantly over subsequent releases. By 2025, Mistral Small 3.1 arrived with multimodal image understanding, multilingual support, and performance benchmarks that surpassed comparable proprietary models, released under an Apache 2.0 open license. The lineage continued with Mistral Small 4, a 119-billion-parameter Mixture of Experts model that activates roughly 6.5 billion parameters per token across 128 experts, consolidating instruction following, reasoning, and coding capabilities into a single unified architecture. This latest iteration also introduced a configurable reasoning mode—allowing fast responses or deeper test-time compute when needed—alongside support for speculative decoding and quantized deployments for performance tuning.