Model details
Nomic Embed Text v1.5
Nomic Embed Text v1.5 is designed to convert text into high-quality embedding vectors for downstream tasks such as semantic search, retrieval-augmented generation, clustering, classification, and data visualization. According to the published model documentation, it is built as a Transformer-based encoder initialized from a BERT-style model (Nomic-BERT-2048) and incorporates rotary embeddings, SwiGLU activations, and long-context adaptations, with an 8192-token context window. The official documentation also states that it is released with open weights, training code, and training data under an Apache-2 license, supporting fully auditable embedding pipelines for both research and enterprise use.
The model uses Matryoshka Representation Learning, which produces resizable embeddings so users can trade off vector size and representation capacity for their retrieval or clustering workload. At inference time, prompts must include task-specific instruction prefixes (such as "search_document:" for documents and "search_query:" for user queries) to guide the embedding behavior in RAG-style applications. A companion vision encoder is aligned to the same embedding space, enabling multimodal retrieval where text embeddings can be matched against image embeddings, while the text-only series remains the supported route for this alignment in current integrations.
Quick Info
Powered by- Provider
- Tinfoil
- Model key
- nomic-embed-text
- Release date
- Feb 1, 2024
- Last updated
- Feb 1, 2024
- Input modalities
- Output modalities
- Capabilities
Cost
A provider subscription or plan supersedes token-based pricing for this model.
Limits
- Output tokens
- 768 tokens
- Context window
- 8,192 tokens
Latest news about Nomic Embed Text v1.5
No articles yet. Fetch the latest news to show it here.