Currently listed through these providers:
Model details
Nvidia Nemotron Nano 12B V2 VL
The Nemotron Nano 12B V2 VL is a multimodal model built on a hybrid Transformer-Mamba architecture, designed to bridge the gap between high-accuracy transformer processing and the memory-efficient sequence modeling of Mamba. This design intent focuses on achieving significantly higher throughput and lower latency, making it a practical choice for complex agentic workflows that require real-time visual and textual reasoning. The model is specifically engineered to handle multi-image document analysis, supporting up to four images at 1k x 2k resolution alongside long text prompts, which allows it to excel in tasks like document summarization, chart reasoning, and visual question answering.
The model lineage is defined by training on high-quality, NVIDIA-curated synthetic datasets that are specifically optimized for optical character recognition and multimodal comprehension. By utilizing Efficient Video Sampling, the model effectively processes long-form video content while maintaining cost-effective inference. Its performance is validated by strong results across benchmarks such as OCRBench v2, MMMU, and MathVista, positioning it as a capable tool for businesses needing to extract and act on information from diverse media like invoices, receipts, and technical manuals. Its architecture and training focus ensure it remains a robust option for developers building next-generation AI agents that require both speed and depth in visual understanding.
Quick Info
Powered by- Provider
- Vercel AI Gateway
- Model key
- nvidia/nemotron-nano-12b-v2-vl
- Release date
- Oct 28, 2025
- Last updated
- Oct 28, 2025
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.20
- Output token cost
- $0.60
Limits
- Output tokens
- 131,072 tokens
- Context window
- 131,072 tokens
Transparent token rates
Compare Nvidia Nemotron Nano 12B V2 VL pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.