Model details
nv-embed-v1
NV-Embed-v1 is a text embedding model engineered to generate high-quality numerical representations from textual inputs, making it well-suited for retrieval-augmented generation workflows and semantic search applications. Its architecture supports a flexible dual-mode operation, where the same model switches between "passage" mode for indexing content into vector databases and "query" mode for encoding search queries, a design choice that significantly impacts retrieval accuracy depending on context. The model handles inputs up to 32k tokens, allowing it to process extended documents or lengthy queries without truncation, and serves as the foundational text embedding backbone for multimodal extensions like MM-Embed, which builds upon its capabilities to support image-text retrieval tasks.
The model demonstrates measurable improvements through its role as the text embedding backbone in MM-Embed, where continual text-to-text fine-tuning lifted its retrieval accuracy from 59.36 to 60.3 across 15 tasks within the Massive Text Embedding Benchmark, while a modality-aware hard negative mining strategy during training further enhanced multimodal retrieval performance, achieving a 52.7 averaged score on the UniIR benchmark compared to the previous best of 48.9. In practical deployments, NV-Embed-v1 integrates into established RAG pipelines alongside frameworks like LangChain, vector databases such as Milvus, and language models including Mistral 7B, providing the semantic understanding layer that bridges user queries with stored knowledge. The open-weights availability enables researchers and developers to fine-tune the model for domain-specific retrieval tasks, while its proven lineage as a foundation for multimodal advances suggests it can serve as a robust base for increasingly sophisticated retrieval systems.
Quick Info
Powered by- Provider
- Nvidia
- Model key
- nvidia/nv-embed-v1
- Release date
- Jun 7, 2024
- Last updated
- Jul 22, 2025
- Input modalities
- Output modalities
- Capabilities
Cost
A provider subscription or plan supersedes token-based pricing for this model.
Limits
- Output tokens
- 2,048 tokens
- Context window
- 32,768 tokens