Currently listed through these providers:
Model details
NVIDIA Nemotron Cascade 2
NVIDIA Nemotron Cascade 2 sits inside the broader Nemotron family of models and is offered as a hosted, open-weights option for developers who want to run NVIDIA-built language models without managing their own infrastructure. Its identity as an openly distributed checkpoint is consistent with the Nemotron line's emphasis on releasing weights so that researchers and practitioners can inspect, fine-tune, and deploy the model on their own terms. The catalog entry indicates that weights are openly available, which makes it a natural fit for experimentation in self-hosted and cloud environments beyond the listed provider.
The model is designed around reasoning, tool calling, and long-context text generation, giving it a profile aimed at agent-style and analytical workloads rather than narrow single-turn chat. The configuration exposed in the catalog reflects a substantial context window paired with a large maximum output length, so it can sustain extended multi-step interactions, retrieval-augmented pipelines, and chained tool invocations within a single conversation. For teams already using NVIDIA hardware or the Nemotron ecosystem, this checkpoint offers a way to add a reasoning-focused, open-weights model to their stack while keeping the deployment choices open.
Quick Info
Powered by- Provider
- Vultr
- Model key
- nvidia/Nemotron-Cascade-2-30B-A3B
- Release date
- Dec 1, 2025
- Last updated
- Dec 1, 2025
- Knowledge cutoff
- 2024-07
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.15
- Output token cost
- $0.60
Limits
- Output tokens
- 131,072 tokens
- Context window
- 262,144 tokens