Sulat.com
AI models
AKI.IO logo

Model details

Mistral Small 4

Mistral Small 4 is positioned as a hybrid general-purpose model that consolidates the roles of three prior families—Instruct, Reasoning, and Devstral—into one unified system, so developers no longer need to pick between separate chat, reasoning, or code-oriented checkpoints for everyday workloads. The model is built on a mixture-of-experts design with 128 experts and 4 active per token, totaling 119B parameters with 6.5B activated per forward pass. This sparse routing is what allows such a large model to behave like a smaller one in latency, and the provider highlights an end-to-end completion time reduction of about 40% in latency-tuned configurations and roughly 3x more requests per second in throughput-tuned configurations compared with the prior generation.

Beyond text, Mistral Small 4 accepts both text and image inputs and produces text outputs, making it suitable for visual question answering, document understanding, and multimodal assistants that previously required an extra vision encoder. It supports a long context window and offers a toggleable reasoning mode that lets callers choose between fast instant replies and deeper, test-time-compute reasoning on a per-request basis, with reasoning effort configurable alongside function calls. Because the weights are openly published, the model can be self-hosted on high-end accelerators, and a community deployment write-up demonstrates running the 119B MoE checkpoint on a single DGX Spark using the SGLang serving stack, pointing to practical viability for local inference at this scale.

AKI.IOmistral4-119bmistral-small

Quick Info

Powered by
Provider
AKI.IO
Model key
mistral4-119b
Release date
Mar 16, 2026
Last updated
Mar 16, 2026
Knowledge cutoff
2025-06
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.20
Output token cost
$0.60

Limits

Output tokens
81,920 tokens
Context window
262,144 tokens

Transparent token rates

Compare Mistral Small 4 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Mistral Small 4

No articles yet. Fetch the latest news to show it here.

Videos about Mistral Small 4

More models around Mistral Small 4