Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
DevPass (LLM Gateway) logo

Model details

Mistral Small 4

The model overview is being prepared.

DevPass (LLM Gateway)mistral-small-2603mistral-small

Quick Info

Powered by
Provider
DevPass (LLM Gateway)
Model key
mistral-small-2603
Release date
Mar 16, 2026
Last updated
Mar 16, 2026
Knowledge cutoff
2025-06
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.15
Output token cost
$0.60

Limits

Output tokens
256,000 tokens
Context window
262,144 tokens

Transparent token rates

Compare Mistral Small 4 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Mistral Small 4

DevPass (LLM Gateway)

Coverage

Mozilla added Mistral Small 4 as a model option in its Firefox Smart Window AI assistant, launching the feature in France with official French-language support. Mozilla cited the model's multilingual performance as a key selection factor, and Smart Window partners including Mistral agree to zero data retention with conversations not saved on Mozilla servers by default. This partnership gives Mistral a consumer route beyond its predominantly enterprise and government sales, while Mozilla frames the deal as a stand against AI gatekeepers and a commitment to open, user-controlled AI in browsing. Smart Window remains optional in the browser, and UK and Germany launches are expected later in 2026 following the France rollout.

DevPass (LLM Gateway)

CoverageBenchmark

Mistral Small 4 launched March 16, 2026 as a 119B-parameter MoE model under Apache 2.0, consolidating Magistral (reasoning), Devstral (coding), and Pixtral (vision) into a single endpoint. The architecture uses 128 expert networks with only 4 active per token, so organizations can collapse three separate deployments, GPU allocations, and monitoring pipelines into one. A reasoning-effort parameter lets developers dial depth per request from Small 3.2-class fast responses up to Magistral-grade step-by-step reasoning, enabling cost-versus-quality optimization without custom routing logic. Coverage also references 40 percent lower latency versus the predecessor and roughly 3x throughput on identical hardware, framed as reducing operational overhead for engineering teams.

DevPass (LLM Gateway)

CoverageBenchmark

A community deployment walkthrough demonstrated running the exact Mistral-Small-4-119B-2603 checkpoint on NVIDIA DGX Spark (GB10/Blackwell) using SGLang with EAGLE speculative decoding. The guide documents NVFP4 quantization at 66 GB on-disk, fitting the model within a single 128 GB unified memory pool without tensor parallelism, while highlighting Blackwell-specific pitfalls such as the sm_121a ptxas crash in the stable SGLang image. Memory analysis shows 29 GB allocated to KV cache at a 65,536 context ceiling, with SGLang allocating 910K total tokens across the pool; the model's native 256K context requires tradeoffs between KV cache capacity and memory. The 390 MB EAGLE draft model adds negligible overhead while reducing latency, making this configuration practical for single-node Blackwell deployment.

DevPass (LLM Gateway)

Coverage

Mistral AI released Mistral Small 4, combining fast text responses, logical reasoning, and image processing in a single model. It has 119B total parameters with only 6B active per query, organized as 128 expert modules with 4 activated at a time. Mistral reports the model is 40 percent faster and handles three times more queries per second than its predecessor. At high reasoning settings, Small 4 matches or exceeds the specialized Magistral line on internal benchmarks, and a configurable control lets users choose between quick or more thorough responses. It ships under Apache 2.0 via Hugging Face, the Mistral API, and Nvidia platforms, and Mistral is joining the Nvidia Nemotron Coalition for open AI model development.

DevPass (LLM Gateway)

CoverageRelease Notes

Mistral released Mistral Small 4 as an open-source model under Apache 2.0, unifying instruct, reasoning, and multimodal capabilities in a single checkpoint. The model uses a Mixture-of-Experts architecture with 128 experts (4 active per token), totaling 119 billion parameters while activating roughly 6 billion per token, with a 256K context window for text and images and configurable reasoning depth. Mistral reports 40% lower latency and 3× higher throughput versus Mistral Small 3. Developers can access Mistral Small 4 through the Mistral API, Mistral AI Studio, Hugging Face, and NVIDIA day-0 NIM containers, with support for major serving frameworks including vLLM and llama.cpp. Enterprises may also deploy it on-premises or fine-tune it for custom needs, underscoring its positioning as a versatile, production-ready foundation model for both research and commercial workflows.

DevPass (LLM Gateway)

CoverageRelease Notes

Mistral AI released Mistral Small 4 on March 16, 2026, as the first model in the Mistral Small line to consolidate instruct, reasoning (Magistral), and multimodal (Pixtral) workloads together with Devstral-style agentic coding into a single deployment target. The model is a Mixture-of-Experts architecture with 128 experts and 4 active per token, totaling 119B parameters with about 6B active (or 8B including embedding/output layers). It supports a 256k-token context window and handles text and image inputs with text output, targeting chat, coding, agentic tasks, and complex reasoning. Mistral positions Small 4 as a general-purpose model that reduces the need to switch models across language-heavy and visually grounded enterprise workflows. The longer context reduces chunking, retrieval orchestration, and pruning overhead in long-document analysis, codebase exploration, multi-file reasoning, and agentic pipelines. The unified design aims to simplify infrastructure by replacing separate reasoning, vision, and coding endpoints with one model.

DevPass (LLM Gateway)

Coverage

Mistral Small 4 launched March 16, 2026 under the Apache 2.0 license, unifying reasoning, vision, and coding into one open-weights model. It uses a Mixture-of-Experts design with 119B total parameters and roughly 6.5B active at inference, paired with a 256K-token context window for processing whole codebases or long documents in a single pass. The release targets hardware-efficient deployment for cost-sensitive enterprise and engineering use cases. A notable feature is a configurable reasoning parameter that lets users adjust computational depth per request instead of routing to a separate reasoning model, eliminating the need for custom switching logic between Magistral, Pixtral, and Devstral endpoints. Multimodal capabilities are integrated natively, so vision tasks no longer require a separate image encoder, which reduces latency and pipeline complexity for teams consolidating AI infrastructure.

Videos about Mistral Small 4

More models around Mistral Small 4