Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
LowRouter logo

Model details

Mistral Medium 3.5

Mistral Medium 3.5 represents a flagship merged release from Mistral AI that consolidates instruction-following, reasoning, and coding capabilities into a single dense 128-billion-parameter model with a 256k context window. Distributed as open weights under a modified MIT license, it replaces Mistral Medium 3.1 and Magistral inside Le Chat, and steps in for Devstral 2 as the engine behind Mistral's Vibe coding agent. The merging strategy allows practitioners to lean on one model across diverse tasks rather than routing between specialized variants, while keeping the weights accessible for self-hosting on modest multi-GPU setups.

Beyond chat, Mistral Medium 3.5 is positioned for long-horizon productivity work: Mistral's remote-agent updates moved coding workflows into the cloud so agents can run asynchronously, and a new Work mode in Le Chat extends the same model to multi-step research, analysis, and tool-driven tasks with function calling and structured output. An NVIDIA NVFP4 quantization broadens deployability for hardware-efficient inference. Practical fit centers on teams and developers who want a unified open-weight model for advanced reasoning, multimodal image understanding, agentic coding flows, and long-context workloads that previously required mixing separate instruct, reasoning, and code models.

LowRouterauto/mistralai/mistral-medium-3.5mistral-medium

Quick Info

Powered by
Provider
LowRouter
Model key
auto/mistralai/mistral-medium-3.5
Release date
Apr 29, 2026
Last updated
Apr 29, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$1.5688
Output token cost
$7.8442

Limits

Output tokens
262,144 tokens
Context window
262,144 tokens

Transparent token rates

Compare Mistral Medium 3.5 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Mistral Medium 3.5

LowRouter

CoverageBenchmark

An NVIDIA DGX Spark user forum post benchmarks Mistral Medium 3.5 128B NVFP4 with the official EAGLE draft model on a single DGX Spark / GB10 node using spark-vllm-docker TF5 and vLLM 0.20.2rc1. Target model memory was 72.91 GiB, KV cache 89,504 tokens, and max concurrency 5.46x at 16k context. The --load-format auto path was stable at 559 seconds load time, while --load-format fastsafetensors triggered OOM at 16k context. Generation throughput at temperature 0.0 ran about 6.9 tok/s for 256-token Q&A, 8.9 tok/s for code and math, and 9.1 tok/s for 2048-token LongCode, all reproduced across two runs. Prompt processing on 3,026 warmup tokens reached 166.2 tok/s, establishing a lower-bound TTFT benchmark for NVFP4 plus EAGLE speculative decoding on consumer-grade DGX Spark silicon.

LowRouter

CoverageDocumentation

NVIDIA NeMo AutoModel documentation covers Mistral Medium 3.5 as a 128B dense flagship that merges instruction-following, reasoning, and coding into one checkpoint with configurable reasoning mode. It unifies the lineage of Mistral Medium 3.1, Magistral Medium, and Devstral 2, and ships natively in FP8 (per-tensor weight scale inv) so the full model fits on a single H200 node or 2× H100 nodes, a footprint advantage over comparable MoE systems. Architecture is Mistral3ForConditionalGeneration (Pixtral vision tower plus dense Ministral-3 text decoder), with 88 decoder layers, hidden 12288, 96 attention heads, 8 KV heads, and GQA, built on the same text backbone as Devstral-2-123B-Instruct-2512. The page documents a 256k context window, 40+ language coverage, and Modified MIT license with a $20M annual revenue threshold, aimed at compact single-node deployment and fine-tuning.

LowRouter

Coverage

The Think Facility model page confirms Mistral Medium 3.5 released on April 28, 2026, with the API model ID mistral-medium-3-5, a 256K token context window, and $1.50 per million input / $7.50 per million output pricing. It notes the model replaces both the Magistral and Devstral product lines, positioning it as Mistral's frontier model with weights released under a modified MIT license. The page provides a pricing and context comparison against GPT-6 Astra, Muse Spark 1.3, Gemini 3.8 Flash, Leanstral 1.5, Mistral Large 3, and Mistral Small 4, placing Mistral Medium 3.5 at $1.50/$7.50 and 256K context. It flags the model as live and currently available, with API integration via mistral-medium-3-5. The page was updated October 10, 2026.

LowRouter

CoverageBenchmark

The AI Release Tracker entry confirms Mistral Medium 3.5 as an open-weight release dated April 29, 2026, 44 days after Mistral Small 4. The model is documented as 128B parameters with a 256k token context window, supports reasoning/thinking, and is priced at $1.50 per million input tokens and $7.50 per million output tokens, with benchmark coverage including SWE-Bench Verified. The tracker also lists Mistral-hosted endpoints including a zdr variant and an EU regional endpoint priced at $1.65/$8.25, all served at 256K context. Rates were verified against mistral.ai on August 18, 2026. The page serves as a release-date and pricing aggregator for tracking Mistral Medium 3.5 availability across regions.

LowRouter

Coverage

Mistral's Hugging Face model card describes Mistral Medium 3.5 128B as the first flagship merged model: a dense 128B parameter checkpoint with a 256k context window that unifies instruction-following, reasoning, and coding in one set of weights. It replaces Mistral Medium 3.1 and Magistral in Le Chat and replaces Devstral 2 in the Vibe coding agent, with reasoning effort configurable per request to switch between quick chat and complex agentic runs. The card documents multimodal input accepting both text and image with text output, plus a vision encoder trained from scratch to handle variable image sizes and aspect ratios. It also notes a companion EAGLE speculative-decoding model for vLLM and SGLang, warns of a Transformers config bug that degraded long-context performance, and ships under a Modified MIT License with revenue-based exceptions.

LowRouter

Official sourceDocumentation

Mistral released Mistral Medium 3.5 as a frontier-class multimodal model optimized for agentic and coding workloads, with open weights under a Modified MIT license. Official docs (v26.04, April 28, 2026) list a 256k context window and pricing of $1.50 per million input tokens and $7.50 per million output tokens, with structured outputs, function calling, predicted outputs, document QnA, prefix caching, and batching supported across the /v1/chat/completions, /v1/agents, and /v1/batch endpoints. The docs page identifies the model as part of the Mistral lineup alongside Z.ai GLM 5.3 and 5.2, and tags it with the Modified MIT license header. It positions Medium 3.5 as a multimodal frontier offering that consolidates prior instruction-following, reasoning, and coding capabilities into a single open-weights release for production agentic use. The page is part of Mistral's inference model catalog served from docs.mistral.ai.

Videos about Mistral Medium 3.5

More models around Mistral Medium 3.5