Azure
Microsoft AI Releases Harrier-OSS-v1: A New Family of Multilingual Embedding Models Hitting SOTA on Multilingual MTEB v2
Model details
text-embedding-3-large is a high-capacity embedding model from OpenAI's text-embedding-3 family, introduced alongside the smaller text-embedding-3-small as part of a January 2024 generation that also refreshed GPT-4 Turbo and GPT-3.5 Turbo. Its core purpose is to convert natural language and code into dense vector representations that downstream systems can use for clustering, semantic search, knowledge retrieval, and retrieval-augmented generation pipelines, including the embeddings that power ChatGPT's knowledge retrieval and the Assistants API.
The model emits 3072-dimensional vectors and is described as the top performer in its family on the MTEB and MIRACL benchmarks, making it well-suited to applications where retrieval quality matters more than compactness. A built-in Matryoshka representation learning capability lets developers truncate vectors to smaller sizes, trading a controllable amount of fidelity for major storage and latency savings in large vector databases. That combination of strong benchmark scores, native dimension reduction, and tight integration with retrieval-oriented workflows makes the model a practical default for production RAG, enterprise search, and recommendation systems that need both accuracy and deployment flexibility.
A provider subscription or plan supersedes token-based pricing for this model.
Azure
Microsoft AI Releases Harrier-OSS-v1: A New Family of Multilingual Embedding Models Hitting SOTA on Multilingual MTEB v2
Azure
Perplexity unveils pplx-embed-v1 and pplx-embed-context-v1, public large-scale retrieval models offering high recall and efficient quantization.
This exact model name is also listed by 3 other providers.