Currently listed through these providers:
Model details
Nemotron-nano 12b v2-vl
The Nemotron Nano 12B v2 VL is built around a hybrid Transformer-Mamba architecture that pairs a CRadioV2-H vision encoder with a 12-billion-parameter language backbone. This combination is engineered to deliver transformer-level accuracy while leveraging Mamba's memory-efficient sequence modeling, which translates into meaningfully higher throughput and lower latency compared to pure transformer designs. The model is architected for document intelligence workloads—capable of processing up to four document images at 1k by 2k resolution alongside a long text prompt, enabling complex multi-image reasoning tasks that go beyond single-image Q&A.
Training leveraged NVIDIA-curated synthetic datasets with particular emphasis on optical character recognition, chart reasoning, and broad multimodal comprehension. The model achieves leading performance on OCRBench v2 and maintains an average score of approximately 74 across a demanding suite including MMMU, MathVista, AI2D, OCRBench, OCR-Reasoning, ChartQA, DocVQA, and Video-MME, surpassing prior open vision-language baselines. Efficient Video Sampling enables cost-effective processing of long-form video content. As an open-weights release under a permissive NVIDIA license, the model, training recipes, and fine-tuning data are available for deployment across NeMo, NIM, and major inference runtimes, making it well-suited for organizations building document processing pipelines or multimodal applications at scale.
Quick Info
Powered by- Provider
- DigitalOcean
- Model key
- nemotron-nano-12b-v2-vl
- Release date
- Oct 28, 2025
- Last updated
- Oct 28, 2025
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.20
- Output token cost
- $0.60
Limits
- Output tokens
- 16,384 tokens
- Context window
- 128,000 tokens