Currently listed through these providers:
Model details
Llama 3.1 Nemotron Nano 8B v1
This compact reasoning model derives from Meta’s Llama 3.1 8B Instruct and is post-trained for conversational work, retrieval-augmented generation, and tool calling. Its target users are teams seeking strong reasoning capability without moving to a much larger model, especially for local deployments and workloads that benefit from efficient inference.
Its post-training combines supervised fine-tuning in mathematics, coding, reasoning, and tool use with reinforcement-learning stages using REINFORCE with RLOO and Online Reward-aware Preference Optimization. The resulting checkpoint aims to improve both reasoning and instruction following, making it a practical fit for local assistants, coding support, retrieval workflows, and other tasks requiring measured reasoning with a long context.
Quick Info
Powered by- Provider
- Nvidia
- Model key
- nvidia/llama-3.1-nemotron-nano-8b-v1
- Release date
- Mar 18, 2025
- Last updated
- Mar 18, 2025
- Input modalities
- Output modalities
- Capabilities
Cost
A provider subscription or plan supersedes token-based pricing for this model.
Limits
- Output tokens
- 16,384 tokens
- Context window
- 131,072 tokens
Latest news about Llama 3.1 Nemotron Nano 8B v1
No articles yet. Fetch the latest news to show it here.