Model details
Titan Text Embeddings V2
Amazon Titan Text Embeddings V2 is a text embedding model purpose-built for retrieval-augmented generation workflows. The model produces fixed-size vector representations of text, with users able to choose between 256-, 512-, or 1024-dimensional outputs depending on their quality and performance trade-offs. It supports over one hundred languages, making it viable for multilingual retrieval systems. The architecture is optimized around two distinct RAG deployment patterns: low-latency single-query embedding via the InvokeModel API for search-time retrieval, and high-throughput batch indexing via Bedrock batch jobs for corpus ingestion.
The model carries MTEB benchmark scores on Hugging Face, providing empirical grounding for retrieval quality claims. Since it returns only embedding vectors, there is no output token charge—users pay only for the embedding generation step. The design reflects a practical split between latency-sensitive and throughput-sensitive retrieval scenarios, allowing teams to match the inference path to their specific pipeline needs. For organizations already invested in the Bedrock ecosystem, Titan Embeddings V2 offers a straightforward way to add semantic search or RAG capabilities without managing a separate embedding service.
Quick Info
Powered by- Provider
- Vercel AI Gateway
- Model key
- amazon/titan-embed-text-v2
- Release date
- Apr 30, 2024
- Last updated
- Apr 1, 2024
- Input modalities
- Output modalities
- Capabilities
Limits
- Output tokens
- 1,536 tokens
- Context window
- 8,192 tokens
Latest news about Titan Text Embeddings V2
No articles yet. Fetch the latest news to show it here.