Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
LowRouter logo

Model details

Mistral Small 4

We haven't written an overview of this model yet. New models can take a few days to gather enough reliable coverage, so check back soon.

LowRouterauto/mistralai/mistral-small-2603mistral-small

Quick Info

Powered by
Provider
LowRouter
Model key
auto/mistralai/mistral-small-2603
Release date
Mar 16, 2026
Last updated
Mar 16, 2026
Knowledge cutoff
2025-06
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.1569
Output token cost
$0.6275

Limits

Output tokens
256,000 tokens
Context window
256,000 tokens

Transparent token rates

Compare Mistral Small 4 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Mistral Small 4

LowRouter

Official sourceAnnouncement

Mistral AI announced Mistral Small 4 on March 16, 2026, unifying the capabilities of Magistral reasoning, Pixtral multimodal, and Devstral agentic coding into a single hybrid model. It is a 119B-parameter Mixture-of-Experts model with 128 experts and 4 active per token, supporting a 256k context window, configurable reasoning effort, and native text-and-image inputs. The model is released under the Apache 2.0 license, and Mistral joined the NVIDIA Nemotron Coalition as a founding member alongside the launch. Performance claims versus Mistral Small 3 include a 40% reduction in end-to-end completion time under a latency-optimized setup and 3x more requests per second in a throughput-optimized setup. Mistral Small 4 activates roughly 6B parameters per token (8B including embeddings), letting users toggle between fast instruct responses and deeper reasoning-intensive outputs on demand. The release positions Small 4 as a single versatile model for chat, coding, agentic tasks, and complex reasoning.

LowRouter

CoverageDocumentation

The NVIDIA NIM catalog page for Mistral-Small-4-119B-2603 confirms the model unifies Instruct, Reasoning (formerly Magistral), and Devstral capabilities into a single hybrid system with multimodal document and image understanding. Deployment is global under the NVIDIA API Trial Terms, with model use governed by the NVIDIA Open Model License; the underlying weights remain Apache 2.0. NVIDIA explicitly notes that the model is not owned or developed by NVIDIA, preserving Mistral attribution. Efficiency tooling on NIM includes speculative decoding via the trained eagle head Mistral-Small-4-119B-2603-eagle and 4-bit float precision via the NVFP4 checkpoint. The page reiterates the 40% latency reduction and 3x throughput gain versus Mistral Small 3, positioning the model for general chat assistants, coding agents, document understanding, and math or research workloads suitable for further customization.

LowRouter

CoverageDocumentation

NVIDIA NeMo AutoModel documents the exact 2603 revision of Mistral Small 4 as a multimodal text+image-to-text model with 119B MoE parameters, architected as MistralForConditionalGeneration under the mistralai Hugging Face organization. The framework page confirms the canonical HF ID mistralai/Mistral-Small-4-119B-2603 for fine-tuning. A supervised fine-tuning recipe on MedPix-VQA is shipped for medical visual question answering workflows. The NeMo AutoModel fine-tuning recipe was validated on 4 nodes x 8 H100 GPUs (32 H100s) for the 2603 variant, providing a concrete multi-node launch baseline. Setup follows the standard NeMo AutoModel container or installation steps, with configuration via mistral4_medpix.yaml. NVIDIA also publishes a broader VLM fine-tuning guide as a related resource for adapting Mistral Small 4.

LowRouter

Coverage

The official Hugging Face model card for mistralai/Mistral-Small-4-119B-2603 details 128 experts with 4 active, 119B total parameters, 6.5B activated per token, a 256k context length, and multimodal text-and-image input with text output. Reasoning mode, function calling, and per-request reasoning effort configuration are all supported alongside a multilingual reach including English, French, Chinese, Japanese, Korean, and Arabic. The artifact ships under the Apache 2.0 license for both commercial and non-commercial use. Two efficiency add-ons are published alongside the base weights: a trained speculative-decoding eagle head at mistralai/Mistral-Small-4-119B-2603-eagle and a 4-bit NVFP4 quantization checkpoint at mistralai/Mistral-Small-4-119B-2603-NVFP4. Recommended settings suggest reasoning effort "high" with temperature 0.7 for complex tasks, and lower temperatures for instant-reply mode. Native JSON output and system-prompt adherence are highlighted as agentic strengths.

Videos about Mistral Small 4

More models around Mistral Small 4