Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
LowRouter logo

Model details

Mistral Large 3

We haven't written an overview of this model yet. New models can take a few days to gather enough reliable coverage, so check back soon.

LowRouterauto/mistralai/mistral-large-2512mistral-large

Quick Info

Powered by
Provider
LowRouter
Model key
auto/mistralai/mistral-large-2512
Release date
Dec 2, 2025
Last updated
Dec 2, 2025
Knowledge cutoff
2024-11
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.5229
Output token cost
$1.5688

Limits

Output tokens
262,144 tokens
Context window
262,144 tokens

Transparent token rates

Compare Mistral Large 3 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Mistral Large 3

LowRouter

Official sourceAnnouncement

Mistral AI announced Mistral 3 on December 2, 2025, headlined by Mistral Large 3, a sparse mixture-of-experts model with 41B active and 675B total parameters, released with base and instruction-tuned weights under Apache 2.0. It is Mistral's first MoE since the Mixtral series, trained from scratch on 3000 NVIDIA H200 GPUs, and debuted at #2 on LMArena among open-source non-reasoning models. The launch confirmed image understanding and best-in-class multilingual conversation performance outside English and Chinese, with a reasoning version to follow. Mistral partnered with NVIDIA, vLLM, and Red Hat to ship an NVFP4 checkpoint built with llm-compressor, enabling efficient runs on Blackwell NVL72 systems or a single 8×A100 / 8×H100 node. All Mistral 3 models were trained on NVIDIA Hopper GPUs.

LowRouter

Coverage

An NVIDIA technical blog dated December 2, 2025 details Mistral 3 deployment on NVIDIA hardware, with the GB200 NVL72 delivering up to 10x the H200 performance for Mistral Large 3, exceeding 5,000,000 tokens per second per megawatt at 40 tokens per second per user. NVFP4 quantization reduces compute and memory costs on the 675B-parameter MoE while using FP8 scaling factors and fine-grained block scaling, and NVIDIA Dynamo disaggregates prefill and decode for long-context inference. Framework support spans TensorRT-LLM, vLLM, SGLang, llama.cpp, and Ollama across everything from GB200 NVL72 datacenters down to RTX AI PCs and Jetson edge devices, with Ministral 3 dense models (3B/8B/14B) shipped in Base, Instruct, and Reasoning variants for nine total checkpoints. The post invites readers to download weights from Hugging Face and try deployment-free evaluation on build.nvidia.com/mistralai.

LowRouter

Coverage

The Hugging Face model card for Mistral-Large-3-675B-Instruct-2512 documents an FP8 post-trained instruct checkpoint built on a granular MoE language model of 673B parameters with 39B active plus a 2.5B vision encoder. It supports a 256k context window, vision input, dozens of languages, system prompts, and native function-calling with JSON output for agentic use. Deployment options listed on the card include FP8 on a single B200 or H200 node, NVFP4 on a single H100 or A100 node, and a BF16 build, all under Apache 2.0. The card names the model Mistral-Large-3-675B-Instruct-2512, explicitly matching the subject alias, and recommends a client-server setup with defined system prompts and sampling parameters.

Videos about Mistral Large 3

More models around Mistral Large 3