Currently listed through these providers:
Model details
Qwen3 Embedding 0.6B
Qwen3 Embedding 0.6B belongs to the Qwen3 Embedding series, a proprietary line from the Qwen family that is specifically designed for text embedding and ranking tasks and built on top of the dense foundational models of the Qwen3 series. The smallest variant in the lineup, it sits alongside 4B and 8B siblings and inherits the multilingual capability, long-text understanding, and reasoning strengths of its underlying Qwen3 base. Packaged on community redistribution hubs, the 0.6B configuration is published in PyTorch and GGUF formats under an Apache-2.0 license, with roughly 595.78M parameters, making it a lightweight option for teams that want to self-host embeddings without the heavier compute footprint of larger encoders.
In practical terms, this model targets the bread-and-butter jobs of any retrieval or analysis pipeline: text and code retrieval, classification, clustering, and bitext mining across many languages. Its longer context behavior, served with a 40K-token window in local runtimes, lets it ingest sizable passages, documents, or mixed-language passages without aggressive chunking, which is useful for semantic search over articles and for cross-lingual matching tasks. For practitioners choosing where Qwen3 Embedding 0.6B fits, the sweet spot is cost-sensitive or latency-sensitive embedding workloads where the additional headroom of the 4B or 8B variants is unnecessary, while still benefiting from the Qwen3 lineage's multilingual coverage and open-weight deployment flexibility.
Quick Info
Powered by- Provider
- NEAR AI Cloud
- Model key
- Qwen/Qwen3-Embedding-0.6B
- Release date
- Jun 3, 2025
- Last updated
- Jun 3, 2025
- Input modalities
- Output modalities
- Capabilities
Cost
A provider subscription or plan supersedes token-based pricing for this model.
Limits
- Output tokens
- 1,024 tokens
- Context window
- 40,960 tokens