Currently listed through these providers:
Model details
Llama 3.3 Nemotron Super 49B v1.5
Llama 3.3 Nemotron Super 49B v1.5 is a reasoning-oriented large language model built by NVIDIA as a derivative of Meta's Llama 3.3 70B Instruct, which serves as the explicit reference model. Rather than training from scratch, the work focused on reshaping the 70B backbone with a novel Neural Architecture Search (NAS) approach that trims the active parameter footprint and memory requirements, allowing the slimmed model to fit on a single high-end GPU such as the H200 while preserving much of the reference's quality. The card frames this as a tunable accuracy versus efficiency tradeoff, positioning v1.5 as a significantly upgraded successor to the earlier Llama 3.3 Nemotron Super 49B v1 checkpoint and aimed at teams that want near-70B reasoning behavior at smaller serving cost.
The model is post-trained specifically for reasoning, conversational quality, and agentic behaviors such as retrieval-augmented generation and tool use, with the final checkpoint produced by merging several reinforcement-learning and preference-optimization stages. It supports a 128K token context window, which is well suited to long-document RAG, multi-turn agent loops, and code or analysis tasks that need to keep substantial material in view. NVIDIA's published accuracy chart on the model card highlights strong benchmark performance relative to peers in its size class, and the practical draw for users is a mid-size reasoning model that can be hosted efficiently while still handling structured reasoning and agentic workflows at scale.
Quick Info
Powered by- Provider
- Deep Infra
- Model key
- nvidia/Llama-3.3-Nemotron-Super-49B-v1.5
- Release date
- Jul 25, 2025
- Last updated
- Jul 25, 2025
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.40
- Output token cost
- $0.40
Limits
- Output tokens
- 131,072 tokens
- Context window
- 131,072 tokens