Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Novita AI logo

Model details

Qwen3 Embedding 8B

We haven't written an overview of this model yet. New models can take a few days to gather enough reliable coverage, so check back soon.

Novita AIqwen/qwen3-embedding-8bqwen

Quick Info

Powered by
Provider
Novita AI
Model key
qwen/qwen3-embedding-8b
Release date
Jun 5, 2025
Last updated
Jun 5, 2025
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
0 tokens
Context window
32,768 tokens

Latest news about Qwen3 Embedding 8B

evroc

CoverageBenchmark

This ResearchGate listing hosts metadata for the Qwen3 Embedding paper (arXiv 2506.05176, June 2025) by Yanzhao Zhang, Mingxin Li, Dingkun Long, Xin Zhang, and collaborators. The abstract directly introduces the Qwen3 Embedding series as an advancement over the GTE-Qwen series, built on Qwen3 foundation models and trai The excerpt explicitly states the series is offered in 0.6B, 4B, and 8B sizes for both embedding and reranking tasks, with the 8B variant serving users who optimize for effectiveness rather than efficiency. Empirical results are reported as state-of-the-art across diverse benchmarks, notably the multilingual MTEB suite

evroc

CoverageBenchmark

Alibaba's Qwen team officially released the Qwen3 Embedding series on June 5, 2025, as a proprietary text embedding and reranking lineup built on the Qwen3 foundation models. The series ships in 0.6B, 4B, and 8B sizes for both embedding and reranking tasks, all open-sourced under Apache 2.0 on Hugging Face and ModelScope, with code and a technical report on GitHub. The 8B embedding variant ranked No. 1 on the MTEB multilingual leaderboard at launch with a score of 70.58. The announcement details strong empirical results across multiple benchmarks, including MTEB-R, CMTEB-R, MMTEB-R, MLDR, MTEB-Code, and FollowIR for the rerankers. It highlights the models' multilingual text understanding, long-context handling, and flexible embedding dimensions that adapt across deployment scenarios. The pipeline combines large-scale unsupervised pre-training, supervised fine-tuning on high-quality datasets, and model merging strategies for robustness.

evroc

CoverageBenchmark

The arXiv technical report 2506.05176 introduces the Qwen3 Embedding series as a major advancement over the GTE-Qwen series, built on Qwen3 foundation models and authored by the Tongyi Lab at Alibaba Group. It describes a multi-stage training pipeline combining large-scale unsupervised pre-training with supervised fine-tuning on high-quality datasets, plus model merging for robustness. The series covers 0.6B, 4B, and 8B sizes for both embedding and reranking tasks, targeting diverse deployment scenarios. The report documents state-of-the-art results across diverse benchmarks, with particular strength on the multilingual MTEB evaluation for text embedding plus code, cross-lingual, and multilingual retrieval. Qwen3 LLMs serve a dual role, acting as backbone models while also synthesizing high-quality, diverse training data across multiple domains and languages. All models are released under Apache 2.0 to support reproducibility and community-driven research.

evroc

Coverage

The official Qwen Hugging Face model card for Qwen3-Embedding-8B lists it as an 8-billion-parameter text embedding model with 32K context length and embedding dimensions up to 4096, supporting user-defined output sizes from 32 to 4096. It documents support for over 100 languages, including programming languages, and strong performance on text retrieval, code retrieval, classification, clustering, and bitext mining tasks. The model ranked No. 1 on the MTEB multilingual leaderboard with a score of 70.58 as of June 5, 2025. The card presents the Qwen3 Embedding series as a cohesive family of text embedding and reranking models in 0.6B, 4B, and 8B sizes, all combining flexibility and multilingual capability. It emphasizes that both embedding and reranking models accept user-defined instructions to boost task-specific performance. Developers are referred to Qwen's blog, GitHub, and the technical report for benchmark evaluation, hardware requirements, and inference performance details.

Videos about Qwen3 Embedding 8B

More models around Qwen3 Embedding 8B