Currently listed through these providers:
Model details
Mistral Small 4
Mistral Small 4 is positioned as a hybrid general-purpose model that consolidates the roles of three prior families—Instruct, Reasoning, and Devstral—into one unified system, so developers no longer need to pick between separate chat, reasoning, or code-oriented checkpoints for everyday workloads. The model is built on a mixture-of-experts design with 128 experts and 4 active per token, totaling 119B parameters with 6.5B activated per forward pass. This sparse routing is what allows such a large model to behave like a smaller one in latency, and the provider highlights an end-to-end completion time reduction of about 40% in latency-tuned configurations and roughly 3x more requests per second in throughput-tuned configurations compared with the prior generation.
Beyond text, Mistral Small 4 accepts both text and image inputs and produces text outputs, making it suitable for visual question answering, document understanding, and multimodal assistants that previously required an extra vision encoder. It supports a long context window and offers a toggleable reasoning mode that lets callers choose between fast instant replies and deeper, test-time-compute reasoning on a per-request basis, with reasoning effort configurable alongside function calls. Because the weights are openly published, the model can be self-hosted on high-end accelerators, and a community deployment write-up demonstrates running the 119B MoE checkpoint on a single DGX Spark using the SGLang serving stack, pointing to practical viability for local inference at this scale.
Quick Info
Powered by- Provider
- AKI.IO
- Model key
- mistral4-119b
- Release date
- Mar 16, 2026
- Last updated
- Mar 16, 2026
- Knowledge cutoff
- 2025-06
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.20
- Output token cost
- $0.60
Limits
- Output tokens
- 81,920 tokens
- Context window
- 262,144 tokens
Transparent token rates
Compare Mistral Small 4 pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about Mistral Small 4
No articles yet. Fetch the latest news to show it here.
Videos about Mistral Small 4
More models around Mistral Small 4
This exact model name is also listed by 11 other providers.