Currently listed through these providers:
Model details
llama-nemotron-rerank-vl-1b-v2
Llama Nemotron Rerank VL 1B v2 is a small vision-language reranker designed to score how well a candidate passage or document matches a text-image query, sitting inside the Nemotron family of specialized retrieval models rather than general chat assistants. Its size class and naming pattern suggest a lightweight transformer that jointly processes text and image inputs to produce a single relevance signal, making it suitable as a second-stage reranker on top of an embedding retriever in multimodal search, image-grounded question answering, and document triage workflows. Because it accepts text plus image inputs but only emits text, it fits cleanly into pipelines that already have a vision encoder upstream and just need a sharp, low-latency relevance decision downstream.
In practical terms, the model is attractive for teams that want a free, openly available reranker with a generous the cataloged API limit token context, large enough to evaluate long passages, multi-image evidence packs, or hybrid text-image candidates in a single pass. Its narrow focus on ranking is reflected in third-party capability scores that credit it strongly on cost efficiency while marking reasoning as unsupported, so it should be paired with a stronger generator for any final answer synthesis. The fit is best for high-volume retrieval-augmented systems, visual document search, and enterprise knowledge bases where ranking quality on text-image matches matters more than open-ended generation.
Quick Info
Powered by- Provider
- Nvidia
- Model key
- nvidia/llama-nemotron-rerank-vl-1b-v2
- Release date
- Mar 31, 2026
- Last updated
- Mar 31, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
A provider subscription or plan supersedes token-based pricing for this model.
Limits
- Output tokens
- 4,096 tokens
- Context window
- 128,000 tokens