Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
LowRouter logo

Model details

Ministral 3 3B

Ministral 3 3B is the smallest and most efficient member of Mistral AI's Ministral 3 family, built for edge deployment while still offering both language and vision capabilities. Released on December 2, 2025, it is distributed under the Apache 2.0 license as an open-weight model, making it accessible for local and on-device setups. Mistral AI positions it for environments that need a compact footprint without sacrificing multimodal coverage, pairing well with the family's broader context handling and structured generation features.

As the direct replacement for the earlier Ministral 3B (v24.1), which Mistral deprecated on the same December 2025 cutover date, Ministral 3 3B carries forward the edge-focused design intent while adding vision support. The launch ties into Mistral AI's Mistral 3 announcement and an associated technical report on arXiv, signaling an updated architecture and training approach for the 3B-scale tier. Practically, it fits lightweight assistants, local inference pipelines, and developer workflows that want open weights, a long context window, and the option to combine text and image inputs in a single small model.

LowRouterauto/mistralai/ministral-3b-2512ministral

Quick Info

Powered by
Provider
LowRouter
Model key
auto/mistralai/ministral-3b-2512
Release date
Dec 2, 2025
Last updated
Dec 2, 2025
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.1046
Output token cost
$0.1046

Limits

Output tokens
262,144 tokens
Context window
262,144 tokens

Transparent token rates

Compare Ministral 3 3B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Ministral 3 3B

LowRouter

Official sourceAnnouncement

Mistral's official announcement of Mistral 3 on December 2, 2025, introduces three small dense models (14B, 8B, and 3B) alongside Mistral Large 3, a sparse mixture-of-experts with 41B active and 675B total parameters. All models are released under Apache 2.0, with the Ministral models positioned as offering the best performance-to-cost ratio in their category. The announcement highlights partnerships with NVIDIA, vLLM, and Red Hat to deliver optimized checkpoints, including an NVFP4 format for efficient inference on Blackwell NVL72 systems and 8×A100 or 8×H100 nodes. Mistral Large 3 debuted at #2 on the LMArena OSS non-reasoning leaderboard, and all Mistral 3 models were trained on NVIDIA Hopper GPUs to maximize accessibility for the open-source community.

LowRouter

CoverageBenchmark

Opper AI's model directory confirms Ministral 3 3B's release on December 2, 2025, with $0.10/M pricing for both input and output tokens. Independent benchmark scores place it in the Efficient tier with global rank 670 of 687 LLMs, showing MMLU-Pro at 52%, GPQA Diamond at 36%, and AIME 2025 at 22%. The page reports an Intelligence Index of 4.8, Coding Index of 4.8, and Math Index of 22.0, with output speed of 206 tokens per second and a 0.57s first-token latency. It notes that Ministral 3 3B is not currently routed through the Opper gateway, but is tracked for its release history and benchmark coverage.

LowRouter

Coverage

The arXiv technical report (2601.08584v1, January 13, 2026) introduces the Ministral 3 series as a family of parameter-efficient dense language models in three sizes: 3B, 8B, and 14B parameters. Each size ships in three variants—base, instruction-tuned, and reasoning—with image understanding capabilities, all under Apache 2.0 license and context lengths up to 256k tokens (128k for reasoning variants). The paper details the Cascade Distillation training recipe, an iterative pruning and continued training with distillation technique derived from the Mistral Small 3.1 24B parent model. Models were trained on between 1 and 3 trillion tokens, achieving competitive performance despite a significantly smaller training budget compared to models like Qwen3 or Llama3, with weights hosted on Hugging Face.

LowRouter

CoverageBenchmark

Artificial Analysis provides an independent evaluation of Ministral 3 3B, confirming its December 2025 release as an open-weight 3B-parameter model under Apache 2.0. The model supports text and image input with text output, has a 256k context window, and offers a 90% cache discount bringing cached input costs to $0.01 per million tokens. Performance metrics show an Intelligence Index of 5, placing it below average among comparable open-weight non-reasoning models, while output speed reaches 230.9 tokens per second, making it notably fast. Pricing at $0.10/M for both input and output is flagged as expensive relative to similar-sized open-weight models, with Hugging Face weights available for self-hosting.

LowRouter

CoverageBenchmark

The AI Release Tracker entry confirms Ministral 3 3B-25-12 as an open-weight model released by Mistral on December 2, 2025, approximately 165 days after Mistral Small 3.2. It lists API pricing at $0.10 per million tokens for both input and output, verified against Mistral's official sources. The tracker provides availability details across Mistral endpoints, including standard and zero-data-retention (zdr) tiers at $0.10/M tokens, and a regional EU endpoint priced at $0.11/M. Each listing shows a 128K context window, which conflicts with Mistral's official 256k specification, suggesting a possible tracker data limitation or regional constraint worth noting.

LowRouter

Official sourceDocumentation

Mistral's official documentation page for Ministral 3 3B (v25.12) confirms a GA release on December 2, 2025, under Apache 2.0, identifying the variant as ministral-3b-2512. The model supports text and vision inputs with a 256k context window and offers structured outputs, function calling, document QnA, prefix caching, chat completions, and batching endpoints (/v1/chat/completions, /v1/conversations, /v1/batch). Pricing is set at $0.10 per million tokens for both input and output. The page positions Ministral 3 3B as the smallest and most efficient member of the Ministral 3 family, designed for edge deployment with robust language and vision capabilities across diverse hardware. It is explicitly framed for local and edge setups while maintaining multimodal support.

Videos about Ministral 3 3B

More models around Ministral 3 3B