Currently listed through these providers:
Model details
Qwen3-Embedding-8B
Qwen3-Embedding-8B is the largest variant in the Qwen3 Embedding series, a family of text embedding and ranking models built on the dense foundation models of the Qwen3 lineup. The series ships in three sizes (0.6B, 4B, and 8B parameters), giving developers a spectrum that trades efficiency against effectiveness. Inheriting Qwen3's multilingual pretraining, the embedding models support over 100 languages, including many programming languages, and are positioned for long-text understanding alongside conventional short-document work.
The intended applications span text retrieval, code retrieval, text classification, text clustering, and bitext mining, with companion reranking models that can be combined with the embeddings for two-stage pipelines. The model card reports that the 8B embedding variant held the top position on the MTEB multilingual leaderboard with a score of 70.58 as of June 5, 2025. Weights are distributed openly through Hugging Face and are mirrored on Ollama, where the 8B tag advertises roughly a 40K-token context window and a multi-gigabyte footprint, making the model a practical fit for retrieval-augmented systems, multilingual semantic search, and code search workloads that benefit from a strong general-purpose embedder.
Quick Info
Powered by- Provider
- InferX
- Model key
- Qwen3-Embedding-8B
- Release date
- Jun 5, 2025
- Last updated
- Jun 5, 2025
- Input modalities
- Output modalities
- Capabilities
Cost
A provider subscription or plan supersedes token-based pricing for this model.
Limits
- Output tokens
- 0 tokens
- Context window
- 32,768 tokens