Sulat.com
AI models
Get $10 off from Venice
Venice AI logo

Model details

Mistral Small 4

Mistral Small 4 is designed as a unified hybrid model that consolidates what were once separate lineages into a single system, combining general instruction capabilities, reasoning features previously branded as Magistral, and the agentic coding strengths of Devstral. Built on a mixture-of-experts architecture with 128 experts and roughly 6.5 billion parameters activated per token out of a total 119 billion, the model is engineered to switch between fast instruction-following and deeper reasoning on a per-request basis, giving developers one checkpoint that can serve many task profiles. Its multimodal front end accepts both text and image inputs while returning text, making it suitable for workflows that mix documents, screenshots, or diagrams with conversational or analytical prompts.

In practice, Mistral Small 4 targets a broad sweet spot of everyday enterprise and developer use, from coding assistance and tool-driven agents to multilingual chat and visual question answering. The model is delivered as an open-weight release with a long context window, and it exposes a configurable reasoning effort knob for tuning latency versus answer quality. Compared with the prior Mistral Small generation, reported gains include a 40 percent reduction in end-to-end completion time in latency-optimized serving and roughly three times the request throughput in throughput-optimized deployments, alongside optional efficiency paths such as speculative decoding with a trained eagle head and NVFP4 quantization. These properties make it a flexible choice for teams that want one generalist system to cover reasoning, vision, and agentic work without juggling multiple specialized models.

Venice AImistral-small-2603mistral-small

Quick Info

Powered by
Provider
Venice AI
Model key
mistral-small-2603
Release date
Mar 16, 2026
Last updated
Jun 11, 2026
Knowledge cutoff
2025-06
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.1875
Output token cost
$0.75

Limits

Output tokens
65,536 tokens
Context window
256,000 tokens

Transparent token rates

Compare mistral-small pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Mistral Small 4

Venice AI

Coverage

Mistral AI has released Mistral Small 4, combining fast text responses, logical reasoning, and image processing in one model.

Venice AI

CoverageBenchmark

See performance metrics across providers for Mistral: Mistral Small 4 - Mistral Small 4 is the next major release in the Mistral Small family, unifying the capabilities of several flagship Mistral models into a single system. It combines strong reasoning from Magistral, multimodal understanding from Pixtral, and agenti

Venice AI

CoverageBenchmark

Mistral Small 4 is the next major release in the Mistral Small family, unifying the capabilities of several flagship Mistral models into a single system. $0.15 per million input tokens, $0.60 per million output tokens. 262,144 token context window. Higher uptime with 2 providers. Includes independent benchmarks from Ar

Videos about Mistral Small 4

More models around Mistral Small 4