OpenAI
Microsoft AI Releases Harrier-OSS-v1: A New Family of Multilingual Embedding Models Hitting SOTA on Multilingual MTEB v2
Model details
The text-embedding-3-small model serves as a highly efficient, performant evolution of the earlier ada embedding architecture. Designed to transform text into numerical representations, it enables systems to measure the relatedness between different pieces of information with precision. This model is built to support a wide array of practical applications, including search, clustering, recommendation engines, anomaly detection, and classification tasks, making it a versatile tool for developers looking to integrate semantic understanding into their workflows.
As a successor in the embedding series, this model emphasizes speed and cost-effectiveness, providing a streamlined solution for large-scale data processing. By utilizing 1536-dimensional embeddings, it balances computational efficiency with the ability to capture complex textual relationships. Its design is particularly well-suited for developers who require a reliable, high-speed foundation for building scalable applications, ensuring that even extensive datasets can be processed effectively without sacrificing the quality of the underlying semantic analysis.
A provider subscription or plan supersedes token-based pricing for this model.
OpenAI
Microsoft AI Releases Harrier-OSS-v1: A New Family of Multilingual Embedding Models Hitting SOTA on Multilingual MTEB v2
OpenAI
Perplexity unveils pplx-embed-v1 and pplx-embed-context-v1, public large-scale retrieval models offering high recall and efficient quantization.
OpenAI
Vercel AI Gateway lists OpenAI's text-embedding-3-small as an actively routed model, confirming $0.02 per 1M input tokens across both Azure and OpenAI providers, with a release date of 01/25/2024. The page documents the model's default 1536 dimensions and an adjustable dimensions parameter for storage optimization, and According to the supplied Vercel page excerpt, text-embedding-3-small delivers higher MTEB scores than its predecessor ada-002 at lower cost, reaching 62.3% on the Massive Text Embedding Benchmark (a 1.3-point gain) and 44.0% on MIRACL for multilingual retrieval (a 12.6-point gain over ada-002's 31.4%). The page is par
This exact model name is also listed by 3 other providers.