Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Vercel AI Gateway logo

Model details

Qwen 3.8 Omni Flash

Qwen 3.8 Omni Flash is Alibaba's native multimodal model designed for audio and video understanding, accepting text, images, audio, and video as input within a single unified interface rather than requiring separate pipelines for each modality. This unified approach lets developers feed a video alongside written directions and have the model identify key moments, streamlining workflows that previously required stitching together multiple specialized systems. The release positions the model as a practical tool for large multimedia analysis tasks where context spans across visual and auditory signals.

Built for long-horizon reasoning over rich media, the model pairs its broad modality support with a one-the cataloged API limit, built-in thinking support, custom function calling, and web search capabilities, enabling it to break down complex multimodal inputs and chain external tools to complete tasks. Alibaba reports over 26% average gains across 30 tests compared with Qwen 3.5-Omni-Plus, reflecting a meaningful step forward in the Qwen omni family. The model is accessible through Alibaba Cloud Model Studio across multiple regions, making it a practical fit for teams building production applications that demand extended context, multimodal reasoning, and tool-augmented generation.

Vercel AI Gatewayalibaba/qwen3.8-omni-flashqwen

Quick Info

Powered by
Provider
Vercel AI Gateway
Model key
alibaba/qwen3.8-omni-flash
Release date
Sep 17, 2026
Last updated
Sep 17, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.15
Output token cost
$0.47

Limits

Output tokens
131,072 tokens
Context window
1,000,000 tokens

Transparent token rates

Compare Qwen 3.8 Omni Flash pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen 3.8 Omni Flash

Vercel AI Gateway

Coverage

This article independently confirms Alibaba's launch of Qwen3.8-Omni-Flash as a native multimodal model that accepts text, images, audio, and video, removing the need for developers to set up separate systems per modality. The release ships with built-in thinking support, custom function calling, and web search, so the The piece highlights the 1M-token context window as the central differentiator, with Alibaba Cloud documentation showing a top input size of 991,808 tokens in non-thinking mode and 983,616 tokens in thinking mode. The authors add a useful caveat that a 1M-token window does not directly translate to unlimited audio or v

Vercel AI Gateway

CoverageRelease Notes

Alibaba's Qwen team has released Qwen3.8-Omni-Flash, described as its first omni-modal model built around agentic capabilities. According to the article, the model accepts text, images, audio, and video inputs and returns text only, combining audio-video understanding, reasoning, and tool use in a single model whose st Technically, Qwen3.8-Omni-Flash is built on the Qwen3.8-Flash-Next architecture (whose open-weights base shipped in August 2026) and offers a 1M-token context window, with QwenCloud listing 991K max input and 131K max output and a max reasoning length of 262K tokens. Thinking mode is on by default with reasoning effort

Vercel AI Gateway

CoverageBenchmark

Qwen3.8-Omni-Flash is Alibaba's next-generation omnimodal model succeeding Qwen3.5-Omni-Plus, built around audio and video agents rather than just multimodal understanding. It processes text, image, audio, and video in one model with a 1M-token context window, while keeping text performance comparable to a text-only mo Headline benchmark gains include a +36.5 point improvement on WildClawBench-MM and a 98%+ reduction in the per-hour price of audio input versus the predecessor. Demonstrated agent capabilities include watching and listening, planning, and calling tools to deliver finished work such as music videos and film commentary.

Vercel AI Gateway

CoverageBenchmark

Alibaba's Qwen team launched Qwen3.8-Omni-Flash on September 18, 2026, positioning it as a shift from understanding omnimodal content to agentic action — planning tasks, calling tools, and completing creative work. The model is a native omnimodal system accepting text, image, audio, and video inputs with a 1-million-to The launch ships with two new open-source components — Qwen-MM-Plugins (agent harness plugins) and Qwen-Live Harness (a real-time interaction runtime) — enabling workflows like Music2MV (generating music videos from song rhythm, mood, and timing), speaker-aware dialogue translation with voice cloning and dubbing, and f

Videos about Qwen 3.8 Omni Flash

More models around Qwen 3.8 Omni Flash