Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Vercel AI Gateway logo

Model details

Embed 5 Fast

The model overview is being prepared.

Vercel AI Gatewaycohere/embed-v5.0-fast

Quick Info

Powered by
Provider
Vercel AI Gateway
Model key
cohere/embed-v5.0-fast
Release date
Sep 30, 2026
Last updated
Sep 30, 2026
Input modalities
Output modalities
Capabilities
Base catalog fields only

Limits

Output tokens
0 tokens
Context window
128,000 tokens

Latest news about Embed 5 Fast

Vercel AI Gateway

CoverageRelease Notes

Cohere released Embed 5 on October 1, 2026, including the Embed 5 Fast tier as a generally available embedding model for enterprise retrieval. Embed 5 Fast targets latency- and cost-sensitive workloads, priced at $0.08 per million tokens, while the Pro tier costs $0.12 per million tokens. Embed 5 Fast and Pro share a single embedding space, so teams can index documents with Pro and query with Fast without rebuilding the index. The family supports multimodal, multilingual, financial, code, and parsed-document retrieval, and is available on the Cohere API and Model Vault, Microsoft Foundry, and Amazon SageMaker.

Vercel AI Gateway

CoverageRelease Notes

Cohere's Embed 5 family lets developers build search indexes with Pro and serve live queries with Fast while staying within a shared vector space. Cohere recommends indexing documents with Embed 5 Pro and routing incoming searches to Embed 5 Fast to favor retrieval quality during preparation and low latency at query time. Embed 5 handles text, images, or mixed inputs including parsed PDF pages, supports more than a hundred languages, and accepts an input window of up to 128,000 tokens. Cohere describes gains over Embed 4 for visually rich documents, financial filings, parsed PDFs, code, and multilingual retrieval, though no quantitative quality-vs-speed benchmark is provided.

Vercel AI Gateway

Coverage

Cohere's Embed 5 launch introduces two tiers: Embed 5 Pro, optimized for retrieval quality, and Embed 5 Fast, designed for latency- and cost-sensitive applications. Both models share a single embedding space, allowing developers to index with Pro and query with Fast without rebuilding existing indexes. Cohere reports an average ViDoRe V3 score of 85.8 for Pro and 84.5 for Fast, indicating a narrow quality gap between the tiers. The family supports multilingual, multimodal, financial, code, and parsed-document retrieval, and is deployable through the Cohere API, Model Vault, Microsoft Foundry, Amazon SageMaker, and private infrastructure.

Vercel AI Gateway

CoverageBenchmark

A third-party benchmark mirror reports Cohere's published document-throughput figures for the Embed 5 launch, with Embed 5 Fast leading the snapshot at 377.3 documents per second. Embed 5 Pro is listed at 159.7 documents per second on the same view. The figures are sourced from a Cohere provider report and hardware, batch size, and serving settings are not disclosed, so the comparison reflects provider throughput rather than end-to-end query latency. BenchLM explicitly excludes these provider comparisons from weighted model rankings.

Videos about Embed 5 Fast