Currently listed through these providers:
Model details
NVIDIA Nemotron Nano 12B v2 VL BF16
The Nemotron Nano 12B v2 VL BF16 is a vision-language model that brings together a C-RADIOv2-H vision encoder and a 12-billion-parameter language backbone derived from the Nemotron-Nano-V2 foundation. This architecture is purpose-built for document-intensive workflows: it can process multiple images at once—up to four at 1k × 2k resolution each—alongside a lengthy text prompt, making it well-suited for handling complex documents such as invoices, receipts, and manuals. The model supports multi-image reasoning, video understanding, visual question answering, and summarization, placing it in the practical domain of multimodal document intelligence rather than open-ended chat.
The model builds on the established Nemotron-Nano-V2 large language model, incorporating a vision encoder designed for efficiency and seamless integration of visual and textual inputs. Optimized for high-throughput inference on NVIDIA hardware, it is engineered to perform well in single-datacenter GPU deployments, which speaks to its practical fit for production workloads rather than research exploration. Its 128,000-token context capacity allows processing of extensive multimodal sequences without losing fidelity, and its performance strength in document intelligence and video comprehension tasks makes it a practical choice for businesses that need to automate processing of mixed media at scale.
Quick Info
Powered by- Provider
- Amazon Bedrock
- Model key
- nvidia.nemotron-nano-12b-v2
- Release date
- Oct 28, 2025
- Last updated
- Oct 28, 2025
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.20
- Output token cost
- $0.60
Limits
- Output tokens
- 8,192 tokens
- Context window
- 131,072 tokens
Latest news about NVIDIA Nemotron Nano 12B v2 VL BF16
No articles yet. Fetch the latest news to show it here.