Currently listed through these providers:
Model details
Llama 3.3 Nemotron Super 49B v1.5
Llama 3.3 Nemotron Super 49B v1.5 is a large language model in NVIDIA's Nemotron family that is explicitly described as a derivative of Meta's Llama 3.3 70B Instruct, distilled or otherwise refined down to a 49 billion parameter footprint for more efficient deployment. NVIDIA positions it as a significantly upgraded successor to the earlier Llama 3.3 Nemotron Super 49B v1, suggesting iterative improvements in alignment, instruction following, or task quality rather than a redesign from scratch. Because it inherits its base from the Llama 3.3 70B instruct lineage, the model is oriented toward general purpose assistant use, including open ended dialogue, reasoning, and tool assisted workflows where a compact yet capable model is preferred over the larger 70B parent.
NVIDIA distributes this version primarily through its NVIDIA NGC catalog as a prebuilt 25h2 container under the NVIDIA NIM and NVIDIA AI Enterprise product lines, framing it as an enterprise ready LLM that can be pulled and served on optimized infrastructure. The same model is also offered through AWS Marketplace, giving organizations an alternative procurement path without changing the underlying weights. Practically, the v1.5 release is aimed at teams that want Llama 3.3 class instruction quality at roughly two thirds the parameter count, making it well suited for latency sensitive or cost constrained deployments while staying inside the established Nemotron Super family.
Quick Info
Powered by- Provider
- Nvidia
- Model key
- nvidia/llama-3.3-nemotron-super-49b-v1.5
- Release date
- Jul 25, 2025
- Last updated
- Jul 25, 2025
- Input modalities
- Output modalities
- Capabilities
Cost
A provider subscription or plan supersedes token-based pricing for this model.
Limits
- Output tokens
- 65,536 tokens
- Context window
- 131,072 tokens
Latest news about Llama 3.3 Nemotron Super 49B v1.5
No articles yet. Fetch the latest news to show it here.