Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Nvidia logo

Model details

Mistral Medium 3.5

Mistral Medium 3.5 is a 128-billion-parameter dense model with a 256k context window, positioned in the mid-tier of the Mistral family for production workloads that require both depth and throughput. Its release is paired with Work Mode in Le Chat, signaling an emphasis on agent-style task automation rather than purely conversational use. Independent commentary highlights the model as a strong fit for asynchronous coding agents, where it can plan, draft, and revise code while retaining long instruction histories within its extended context.

Practically, the model accepts both text and image inputs while producing text outputs, allowing it to serve multimodal pipelines such as document analysis, UI reasoning, and code-in-image scenarios. Because the weights are open, teams can self-host on appropriately scaled hardware, and community benchmarks on single-node DGX Spark setups have explored NVFP4 quantization combined with EAGLE speculative decoding to push token throughput on the 128B configuration. The combination of long context, open availability, and coding-agent orientation makes it well suited for organizations building private agent stacks or replacing closed mid-size models with a self-managed alternative.

Nvidiamistralai/mistral-medium-3.5-128bmistral-mediumdeprecated

Quick Info

Powered by
Provider
Nvidia
Model key
mistralai/mistral-medium-3.5-128b
Release date
Apr 29, 2026
Last updated
Apr 29, 2026
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
32,768 tokens
Context window
262,144 tokens

Latest news about Mistral Medium 3.5

Nvidia

Official sourceDocumentation

NVIDIA NIM for Vision Language Models version 2.1.1 lists Mistral Medium 3.5 as a supported model, alongside a separate "Mistral Medium 3.5 128B (NIM Certified)" variant. The catalog page documents 19 versioned releases spanning from 1.0.0 through 2.1.2-variant, confirming active multi-version NIM support for this mode The NIM for VLMs overview positions the platform as bringing state-of-the-art VLMs to enterprise self-hosted environments, with industry-standard APIs for building copilots, chatbots, and AI assistants, leveraging NVIDIA GPU acceleration. Public-catalog NIMs can be pulled and run without an NGC API key, while certified

Nvidia

Coverage

The news blog specialized in Japanese culture, odd news, gadgets and all other funny stuffs. Updated everyday.

Nvidia

Coverage

Mistral AI unveils Mistral Medium 3.5, a 128B dense model with 256k context, offering global access, cloud agents, and workflow automation.

Nvidia

CoverageRelease Notes

The official Mistral changelog records that on April 27, 2026, Mistral released Mistral Medium 3.5 under the model identifier mistral-medium-3-5, alongside separate entries for OCR 4.1 going Generally Available, OCR 4 release, and the Leanstral 1.5 Lean 4 formal proof engineering model. The same changelog separately do OCR-related API updates in the same period include the OCR API confidence scores granularity parameter now supporting "block" granularity (with "page" returning page-level scores only and "word" returning page-level and word-level scores), the introduction of include blocks that return a blocks array with paragraph-lev

Nvidia

Coverage

The ThursdAI aggregator's Mistral AI releases index lists Medium 3.5 under April 30, 2026, confirming it as a 128B dense flagship model with 256K context and configurable reasoning, released with weights on Hugging Face and accompanied by a Vibe coding agent launch. The entry points to Mistral's blog, the Hugging Face The broader index catalogs 15 Mistral AI releases covered on the podcast between January 2025 and July 2026 — including Robostral Navigate (July 8, 2026, an 8B embodied-navigation model claiming SOTA on R2R-CE), Mistral Small 4 (March 19, 2026), Voxtral TTS (March 26, 2026, a 3B open-weight text-to-speech model), Mistr

Nvidia

CoverageBenchmark

OpenRouter's listing independently corroborates Mistral Medium 3.5's core specs: 128B dense, 262K context (noting 256K usable plus overhead), text-and-image input with text output, configurable reasoning effort per request, and a custom vision encoder handling variable image sizes. Listed pricing is $1.50 per million i OpenRouter exposes three Mistral routing variants for this model — Mistral, Mistral (EU), and Mistral (ZDR / Zero Data Retention) — with the EU endpoint priced slightly higher at $1.65/$8.25 per million tokens. Observed P50 latency is 0.39s on Mistral and Mistral ZDR (0.83s on EU), with throughput up to 161 tok/s P50 o

Videos about Mistral Medium 3.5

More models around Mistral Medium 3.5