Model details
Mistral-Nemo-Instruct-2407
Mistral-Nemo-Instruct-2407 is a transformer-based model built with a 12-billion parameter architecture that serves as a robust successor to smaller 7B-class models. Designed with 40 layers, a 5,120-dimension hidden state, and a 128k context window, the model utilizes SwiGLU activation and grouped-query attention to balance efficiency with deep comprehension. Its design intent centers on providing a highly capable, drop-in replacement for smaller architectures, offering enhanced world knowledge and improved performance across a wide range of conversational, instructional, and coding-heavy applications.
Developed through a joint collaboration between Mistral AI and NVIDIA, the model underwent extensive instruction tuning to refine its ability to follow complex prompts and handle multilingual data. This training lineage enables the model to demonstrate strong performance across diverse languages, including French, German, Spanish, and Japanese, while maintaining high accuracy on benchmarks like MMLU and TriviaQA. With its native support for function calling and structured output, the model is well-positioned for developers building agentic workflows and interactive applications that require reliable, context-aware responses.
Quick Info
Powered by- Provider
- OVHcloud AI Endpoints
- Model key
- mistral-nemo-instruct-2407
- Release date
- Nov 20, 2024
- Last updated
- Nov 20, 2024
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.14
- Output token cost
- $0.14
Limits
- Output tokens
- 65,536 tokens
- Context window
- 65,536 tokens