Sulat.com
AI models
Meganova logo

Model details

Mistral Nemo Instruct 2407

Mistral Nemo Instruct 2407 is an instruct fine-tuned version of Mistral Nemo Base 2407, developed jointly by Mistral AI and NVIDIA and released under the Apache 2 license. It is designed as a drop-in replacement for Mistral 7B, with a transformer architecture comprising 40 layers, a hidden dimension of 5,120, a SwiGLU activation function, 32 attention heads paired with 8 key-value heads under grouped-query attention, and a vocabulary of roughly 128k tokens using rotary embeddings scaled to theta of 1M. The model was trained with a 128k context window on a substantial share of multilingual and code data, aiming to push performance beyond earlier models of comparable size while remaining open-weight for research and product use.

On common reasoning and knowledge benchmarks the model reports competitive results for its class, including 68.0% on MMLU 5-shot, 83.5% on HellaSwag 0-shot, and 76.8% on Winogrande 0-shot, alongside balanced multilingual MMLU scores across French, German, Spanish, Italian, Portuguese, Russian, and Chinese in the low-to-mid 60s and 59% range for Japanese. These figures suggest practical strength for conversational assistants, multilingual chat, and code-aware workflows that benefit from a very long context. The release is suited to teams that want an open multilingual model with broad tokenizer coverage, long-context handling, and flexible deployment through frameworks such as mistral-inference or transformers.

Meganovamistralai/Mistral-Nemo-Instruct-2407mistral

Quick Info

Powered by
Provider
Meganova
Model key
mistralai/Mistral-Nemo-Instruct-2407
Release date
Jul 18, 2024
Last updated
Jul 18, 2024
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.02
Output token cost
$0.04

Limits

Output tokens
65,536 tokens
Context window
131,072 tokens

Transparent token rates

Compare mistral pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Mistral Nemo Instruct 2407

Meganova

CoverageBenchmark

Compare Grok 4 Fast vs Mistral NeMo Instruct: input $0.2/M vs $0.15/M, output $0.5/M vs $0.15/M tokens. Mistral NeMo Instruct is 133% cheaper overall. Full API cost breakdown, context window, and benchmark comparison.

Videos about Mistral Nemo Instruct 2407

More models around Mistral Nemo Instruct 2407