Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
EmpirioLabs AI logo

Model details

Qwen3.8 Omni Flash

Qwen3.8 Omni Flash appears in the Qwen model family as an omnimodal entry that has so far been documented only through community and tracker surfaces. A Hacker News discussion titled "Qwen 3.8 Omni Flash" links to a Qwen blog page at qwen.ai/blog?id=qwen3.8-omni-flash, signaling that an announcement post exists on the official Qwen domain even though its contents have not been independently fetched. The model name follows the family lineage associated with Alibaba's Qwen releases, with one tracker page listing the creator attribution as Alibaba in its social card metadata, which is consistent with prior Qwen development rather than with any unrelated provider label.

Early third-party benchmark aggregation suggests the model is positioned as a balanced, general-purpose omnimodal option rather than a single-task specialist. The BenchLM tracker records an "owner-defined voice evaluation" alongside a voice-benchmark subpage, indicating that speech-style inputs and outputs are part of the intended evaluation surface, while overall instruction-following rankings place it as a well-rounded choice across a range of tasks. This combination of omnimodal voice evidence and broadly distributed benchmark rows points to a model designed for mixed text, image, audio, and video interactions, with practical fit for users who want one model for diverse workloads rather than a narrow expert system.

EmpirioLabs AIqwen3-8-omni-flashqwen

Quick Info

Powered by
Provider
EmpirioLabs AI
Model key
qwen3-8-omni-flash
Release date
Sep 17, 2026
Last updated
Sep 17, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.30
Output token cost
$0.94

Limits

Output tokens
131,072 tokens
Context window
1,000,000 tokens

Transparent token rates

Compare Qwen3.8 Omni Flash pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3.8 Omni Flash

EmpirioLabs AI

CoverageBenchmark

The Decoder reports Qwen3.8-Omni-Flash is Qwen's first multimodal model built for AI agents, processing audio and video together to draw conclusions and use tools autonomously for tasks like editing vlogs, translating short videos, or summarizing movies, with a one-million-token context window. Alibaba positions the mo For comparison, Gemini 3.8 Flash's introductory rate is $0.75 input and $3.75 output per million tokens, with prices set to double on January 1, 2027. Access is through Qwen Studio, Qwen Cloud, and the API, with the open-source Qwen-MM-Plugins adding video editing, speaker recognition, PDF video notes, and reusable wor

EmpirioLabs AI

CoverageRelease Notes

MarkTechPost details that Qwen3.8-Omni-Flash is built on the Qwen3.8-Flash-Next architecture, whose base model shipped with open weights in August 2026. The model accepts text, images, audio, and video and returns text only, with QwenCloud listing 991K max input and 131K max output tokens and a maximum reasoning length of 262K tokens. Thinking mode is enabled by default with reasoning effort set to xhigh. Deployment is available as a hosted API on QwenCloud, Alibaba Cloud Model Studio, and Qwen Studio at launch, with no open weights announced, so self-hosting is not yet an option. Developers needing speech generation are directed to Qwen3.5-Omni, since Qwen3.8-Omni-Flash outputs text only, while supporting function calling and web search for agentic workflows.

EmpirioLabs AI

CoverageBenchmark

GIGAZINE reports Alibaba added Qwen3.8-Omni-Flash, a next-generation native omnimodal model with a one-million-token context window, to the Qwen lineup, claiming strong results across workflows including video editing, music video production, film production, audiovisual summarization, and real-time conversation. Acros Specific benchmark numbers reported include 71.0 on WildClawBench-MM (a 36.5-point jump over Qwen3.5-Omni-Plus) and 69.6 on UniClawBench, plus 82.7 on LongAudioSpan, 63.4 on OmniVideoBench, 28.2 on OmniCap-IF, and 89.7 on AliMeeting for Chinese multi-speaker conference transcription. Alibaba says Qwen3.8-Omni-Flash sur

EmpirioLabs AI

Coverage

Alibaba's Qwen team released Qwen3.8-Omni-Flash, a native omnimodal model that accepts text, images, audio, and video inputs and ships with a one-million-token context window. According to Neowin, the model is available through the Qwen chat interface and mobile app (selectable from the model picker) and is reachable t On benchmarks, Qwen3.8-Omni-Flash lifts its OmniVideoBench score from 63.4 to 67.8 while cutting token consumption 45.7% (from 145,736 to 79,117 tokens), reflecting an agentic evidence-gathering approach to long-form video rather than brute-force processing. Alibaba frames the upgrade around controllable descriptions,

EmpirioLabs AI

Coverage

Alibaba's Qwen team released Qwen3.8-Omni-Flash on September 18, 2026, as a next-generation native omnimodal model aimed at shifting from passive understanding to agentic task planning and tool execution in production workflows such as video editing, music video creation, film production, audiovisual summarization, and real-time conversation. The model accepts text, images, audio, and video inputs within a 1M-token context window while keeping text performance comparable to same-size text-only models. The official Qwen blog reports average gains exceeding 25% across 29 evaluations versus Qwen3.5-Omni-Plus, with WildClawBench-MM up 36.5 points, UniClawBench at 69.6, and AgenticVBench up 22.3 points, alongside AliMeeting diarization error rates falling from 88.11 to 3.35. API pricing drops by over 98% per hour for audio input and over 93% for audio-visual input, with overall audio performance exceeding Gemini 3.8 Flash.

EmpirioLabs AI

CoverageBenchmark

GIGAZINE's English coverage confirms Alibaba added Qwen3.8-Omni-Flash to its Qwen series as a next-generation native omnimodal model delivering results across video editing, music video production, film production, audiovisual summarization, and real-time conversation. The model pairs a 1M-token context window with significantly improved omnimodal capability while maintaining text performance equivalent to text-only models of the same size. Reported benchmark scores include LongAudioSpan 82.7, OmniVideoBench 63.4, OmniCap-IF 28.2, and AliMeeting 89.7, with Qwen3.8-Omni-Flash surpassing Gemini 3.8 Flash on multiple tests. API cost per hour of voice input is reported down over 98% and per-hour audio-video input down over 93%, while the model scored 71.0 on WildClawBench-MM, a 36.5-point jump over Qwen3.5-Omni-Plus.

EmpirioLabs AI

Coverage

AIBASE reports that Qwen3.8-Omni-Flash handles text, images, audio, and video natively at the model level rather than by combining separate speech and image recognizers, with the stated goal of enabling the model to understand audio/video and then plan tasks, call tools, and complete work. Compared with Qwen3.5-Omni-Plus, official claims cite average gains above 25% across 29 benchmark tests. On multimodal tool calling the release beats Gemini3.8Flash on WildClawBench-MM (71.0 vs 58.9), DailyOmni (85.1 vs 84.0), SpotSoundBench (67.2 vs 39.7), and MMAU (81.8 vs 76.9), though Gemini still leads on AgenticVBench and some video tests. Alibaba Cloud's international region pricing is $0.15 per million input tokens, $0.016 per million cached input tokens, and $0.47 per million output tokens.

Videos about Qwen3.8 Omni Flash

More models around Qwen3.8 Omni Flash