Currently listed through:
Model details
Mistral Nemo
Mistral Nemo is a 12-billion-parameter language model born from a collaboration between Mistral AI and NVIDIA, engineered to deliver frontier-level intelligence within a compact footprint. Its foundation rests on a standard LLaMA architecture, making it a straightforward drop-in replacement for systems already running Mistral 7B. The model introduces a custom Tekken tokenizer that brings notable efficiency gains for European languages, complementing its broad multilingual capabilities that span English, French, German, Spanish, Italian, Portuguese, Chinese, Japanese, Korean, Arabic, and Hindi. Built with quantization awareness from the ground up, it enables FP8 inference without sacrificing performance, making it particularly well-suited for conversational agents, virtual assistants, and enterprise chatbot deployments that demand speed alongside quality.
The Mistral Nemo Instruct 2407 variant underwent instruction fine-tuning through NVIDIA NeMo, sharpening its ability to follow precise instructions, maintain multi-turn conversations, and handle coding tasks with greater reliability than its predecessors. Training took place on NVIDIA DGX Cloud infrastructure, leveraging TensorRT-LLM for accelerated inference and the NeMo development platform for optimization. The model was developed to target function calling and multilingual knowledge retrieval, positioning it as a practical choice for organizations building global AI applications. Released under Apache 2.0 licensing for both base and instruction-tuned checkpoints, it offers researchers and enterprises an accessible path to customization while delivering state-of-the-art reasoning and world knowledge within its size class.
Quick Info
Powered by- Provider
- Vercel AI Gateway
- Model key
- mistral/mistral-nemo
- Release date
- Jul 18, 2024
- Last updated
- Jul 1, 2024
- Knowledge cutoff
- 2024-04
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.15
- Output token cost
- $0.15
Limits
- Output tokens
- 128,000 tokens
- Context window
- 128,000 tokens
Transparent token rates
Compare mistral-nemo pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about Mistral Nemo
Videos about Mistral Nemo
More models around Mistral Nemo
This exact model name is also listed by 7 other providers.