Sulat.com
AI models
STACKIT logo

Model details

Qwen3-VL Embedding 8B

Built on the open Qwen3-VL foundation, this 8B-parameter embedding model is part of a broader multimodal retrieval suite that also includes a companion reranker for refining search results. It is designed to encode inputs that mix text, images, screenshots, and video into high-dimensional vectors, producing semantically rich representations in a shared visual-textual space. This unified representation makes it well suited to tasks such as image-text retrieval, visual question answering, video-text matching, and multimodal content clustering.

The model supports more than 30 languages and operates within a 32,768-token context window, making it practical for multilingual enterprise search and document discovery workflows. Typical use cases include screenshot search, PDF document retrieval, and any application that needs similarity scoring across both visual and textual content. Because the reranker accepts query-document pairs containing arbitrary single or mixed modalities, teams can pair the embedding model with its reranker counterpart to build a more accurate two-stage multimodal search pipeline without juggling separate vision and language encoders.

STACKITQwen/Qwen3-VL-Embedding-8Bqwen

Quick Info

Powered by
Provider
STACKIT
Model key
Qwen/Qwen3-VL-Embedding-8B
Release date
Feb 5, 2026
Last updated
Feb 5, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.09
Output token cost
$0.09

Limits

Output tokens
4,096 tokens
Context window
32,000 tokens

Latest news about Qwen3-VL Embedding 8B

STACKIT

Official sourceDocumentation

STACKIT's official "Available Shared Models" documentation page (last updated Aug 24, 2026) lists Qwen3-VL-Embedding-8B within its shared AI Model Serving portfolio, served via the OpenAI-compatible base URL https://api.openai-compat.model-serving.eu01.onstackit.cloud/v1. The catalog page is the canonical first-party r Within that same shared-models catalog, the page also details Qwen3-VL 235B (235B parameters in 8-bit quantization, 200K context length, text+image input modalities, tool calling enabled, 350K TPM / 30 RPM limits) and Qwen3.6 27B (coding-focused LLM, 262K context window), giving developers immediate context on the surr

STACKIT

CoverageRelease Notes

OpenVINO 2026.3 release notes (Aug 4, 2026) name Qwen3-VL-Embedding-8B as one of the "additional models available as early releases on CPU and GPU," alongside Qwen3-ASR, Qwen3-Omni, Gemma-3n, and Kokoro-82M. This confirms the model is a deployable multimodal embedder with active ecosystem tooling and that it is recogni The same OpenVINO release adds a new GenAI "Embedding" pipeline for multimodal embedding generation and extends EAGLE-3 speculative decoding, plus NNCF FP8 quantization for ONNX, Hugging Face Transformers v5.5 compatibility, and Intel Xeon 6+ support. For STACKIT customers evaluating on-prem parity or local inference,

STACKIT

CoverageBenchmark

An arXiv preprint titled "The Commercial Tax: Rent-vs-Own Blind Spots in Multi-Hop Retrieval Benchmarks" (Sanchez & Dehnad, 17 Aug 2026) explicitly benchmarks Qwen3-VL-Embedding-8B among thirteen embedders from eight makers on the MuSiQue multi-hop retrieval harness. The model is reported at -9.67 Recall@5 versus the N The same paper notes that NVIDIA's Nemotron-3-Embed-8B has effectively closed the "commercial tax" gap as of mid-2026, while every other commercially-licensed entrant — including Qwen3-VL-Embedding-8B — remains measurably behind the non-commercial anchor on multi-hop retrieval. This independent third-party result gives

Videos about Qwen3-VL Embedding 8B

More models around Qwen3-VL Embedding 8B