Mistral
Mistral NeMo: our new best small model. A state-of-the-art 12B model with 128k context length, built in collaboration with NVIDIA, and released under the Apache 2.0 license.
Model details
Mistral NeMo is a compact 12-billion-parameter language model co-developed with Mistral AI and NVIDIA, designed as a small but capable alternative to much larger frontier systems. The model is built on a LLaMA-style architecture and ships under the permissive Apache 2.0 license, with an unusually long context window that supports extended documents, multi-turn conversations, and long-running agent workflows. A defining piece of its design is the Tekken tokenizer, which was engineered for stronger efficiency across European languages, reinforcing NeMo's positioning as a multilingual-first model. The weights are open, and quantization-friendly variants make it practical to run on modest hardware while still leaving room for customization, which is why it is often framed as an enterprise-friendly foundation that can be fine-tuned, distilled, or deployed behind private infrastructure. The instruct-tuned release, Mistral-Nemo-Instruct-2407, was fine-tuned in collaboration with NVIDIA using the NeMo toolkit to sharpen instruction following, multi-turn dialogue, code generation, and reasoning compared to the earlier Mistral 7B family. That post-training focus is reflected in its intended uses, spanning conversational assistants, knowledge retrieval, and code assistance across English, French, German, Spanish, Italian, Portuguese, Russian, Chinese, and Japanese, with tool-calling support for agentic applications. Its blend of a generous context window, open weights, and a relatively small parameter count makes it a strong fit for on-device or private-cloud deployments where latency, cost, and data control matter, while still delivering the multilingual fluency and instruction behavior that modern assistants require.
Mistral NeMo is a compact 12-billion-parameter language model co-developed with Mistral AI and NVIDIA, designed as a small but capable alternative to much larger frontier systems. The model is built on a LLaMA-style architecture and ships under the permissive Apache 2.0 license, with an unusually long context window that supports extended documents, multi-turn conversations, and long-running agent workflows. A defining piece of its design is the Tekken tokenizer, which was engineered for stronger efficiency across European languages, reinforcing NeMo's positioning as a multilingual-first model. The weights are open, and quantization-friendly variants make it practical to run on modest hardware while still leaving room for customization, which is why it is often framed as an enterprise-friendly foundation that can be fine-tuned, distilled, or deployed behind private infrastructure. The instruct-tuned release, Mistral-Nemo-Instruct-2407, was fine-tuned in collaboration with NVIDIA using the NeMo toolkit to sharpen instruction following, multi-turn dialogue, code generation, and reasoning compared to the earlier Mistral 7B family. That post-training focus is reflected in its intended uses, spanning conversational assistants, knowledge retrieval, and code assistance across English, French, German, Spanish, Italian, Portuguese, Russian, Chinese, and Japanese, with tool-calling support for agentic applications. Its blend of a generous context window, open weights, and a relatively small parameter count makes it a strong fit for on-device or private-cloud deployments where latency, cost, and data control matter, while still delivering the multilingual fluency and instruction behavior that modern assistants require.
Mistral
Mistral NeMo: our new best small model. A state-of-the-art 12B model with 128k context length, built in collaboration with NVIDIA, and released under the Apache 2.0 license.
Mistral
Mistral AI and NVIDIA today released a new state-of-the-art language model, Mistral NeMo 12B, that developers can easily customize and deploy for enterprise...
This exact model name is also listed by 7 other providers.