Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
AIHubMix logo

Model details

Qwen3.8 Omni Flash

Qwen3.8 Omni Flash is presented as the latest omnimodal entry in the Qwen family, surfacing through a Qwen team blog post on qwen.ai that drew notable discussion on Hacker News. According to a third-party explainer, the model is built around a hybrid core that combines Gated DeltaNet with Qwen Sparse Attention, replacing the static perception patterns used in earlier multimodal systems with an agentic perception approach designed to actively query and reason over incoming audio, image, and video streams. That pairing of selective sparse attention with a gated recurrent-style memory component is positioned as the architectural backbone for handling long, mixed-modality sessions while keeping reasoning efficient.

In practical terms, the variant is aimed at developers building assistants that need to interpret live audio and video alongside text, rather than at teams looking only for a pure language model. Third-party coverage highlights native audio-video reasoning, multi-speaker and spatial audio understanding, adjustable reasoning effort, function calling, and an extensible plugin layer called Qwen-MM-Plugins as the main capability surfaces. Together those traits make it a fit for agentic applications such as meeting copilots, surveillance-style scene understanding, and tool-using assistants that must combine speech, visual context, and external actions within a single workflow.

AIHubMixqwen3.8-omni-flashqwen

Quick Info

Powered by
Provider
AIHubMix
Model key
qwen3.8-omni-flash
Release date
Sep 17, 2026
Last updated
Sep 17, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.1126
Output token cost
$0.380025

Limits

Output tokens
131,072 tokens
Context window
1,000,000 tokens

Transparent token rates

Compare Qwen3.8 Omni Flash pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3.8 Omni Flash

AIHubMix

CoverageBenchmark

Qwen3.8-Omni-Flash undercuts Google's Gemini Flash pricing while matching its multimodal benchmarks, according to coverage of Qwen's launch. The model processes audio and video together, draws conclusions, and uses tools on its own to edit vlogs, translate short videos, or summarize movies, with a context window spanni Qwen API pricing sits at $0.15 per million input tokens and $0.47 per million output tokens. Qwen estimates audio input at under $0.01 per hour, while 720p video with audio at one frame per second runs about $0.20, excluding response costs. For comparison, Gemini 3.8 Flash charges $0.75 for input and $3.75 for output p

AIHubMix

CoverageBenchmark

DataCamp's coverage of Qwen3.8-Omni-Flash positions it as Alibaba's next-generation omnimodal model replacing Qwen3.5-Omni-Plus, with the differentiator being audio and video agents: watching, listening, planning, and calling tools to deliver finished work like music videos and film commentary. It accepts text, image, Headline agent gains include 36.5 points on WildClawBench-MM and a reported 98%+ drop in the per-hour price of audio input. At roughly $0.15 input and $0.47 output per 1M tokens, it undercuts most Flash-tier rivals on text pricing. DataCamp recommends Qwen3.8-Omni-Flash for audio- or video-heavy workloads, while noting

AIHubMix

CoverageRelease Notes

MarkTechPost reports that Qwen3.8-Omni-Flash is built on the Qwen3.8-Flash-Next architecture and is live as a hosted API on QwenCloud, Alibaba Cloud Model Studio, and Qwen Studio, with no open weights announced at launch. Thinking mode is enabled by default with reasoning effort set to xhigh, and QwenCloud lists 991K max input, 131K max output, and 262K max reasoning length. The article frames the release as Alibaba's first omni-modal model centered on agentic capabilities, combining audio-video understanding, reasoning, and tool use inside a single model. The stated workflow is understand, plan, execute with tools, and deliver, targeting production scenarios such as video editing, meeting transcription, and real-time conversation.

AIHubMix

CoverageBenchmark

Alibaba's Qwen team launched Qwen3.8-Omni-Flash on September 18, 2026, framing it as a shift "from understanding omnimodal content to planning tasks, calling tools, and completing creative work." It is a 1-million-token-context omnimodal model with audio API pricing down more than 98% versus its predecessor, plus a new A companion realtime variant, Qwen3.8-Omni-Flash-Realtime, targets live speaking practice, spatial audio, and dynamic skills. Demonstrated workflows include Music2MV (plans music video characters, scenes, and shots from rhythm, mood, and lyric timing), a translation pipeline that performs speaker-aware dialogue recogni

AIHubMix

CoverageBenchmark

Qwen3.8-Omni-Flash is Alibaba's next-generation native omnimodal model in the Qwen series, supporting a 1-million-token context window and significantly improved omnimodal functionality while maintaining text performance equivalent to text-only models of the same size. Across 29 benchmarks, the average score is more th On audio and video agent coding and long-term tasks, Qwen3.8-Omni-Flash scored 71.0 on WildClawBench-MM (a 36.5-point gain over Qwen3.5-Omni-Plus) and 69.6 on UniClawBench. It also posted 82.7 on LongAudioSpan, 63.4 on OmniVideoBench, 28.2 on OmniCap-IF, and 89.7 on AliMeeting. Alibaba stated the model surpassed Gemini

AIHubMix

Coverage

Alibaba's Qwen team launched Qwen3.8-Omni-Flash on September 18, 2026, positioning it as a next-generation native omnimodal model built around agentic delivery rather than passive perception. The model accepts text, image, audio, and video inputs with a 1 million token context window, and output is text only, with developers directed to Qwen3.5-Omni for speech generation. Across 29 benchmarks, Qwen3.8-Omni-Flash averages more than 25 percent improvement over Qwen3.5-Omni-Plus, including +36.5 on WildClawBench-MM and 69.6 on UniClawBench for audio-visual agent tasks. Alibaba reports API pricing drops of over 98 percent per hour for audio input and over 93 percent for audio-visual input, while AliMeeting diarization error rates fell from 88.11 to 3.35 DER.

AIHubMix

CoverageBenchmark

GIGAZINE's coverage confirms Qwen3.8-Omni-Flash targets workflows including video editing, music video production, film production, audiovisual summarization, and real-time conversation. The 1M token context window preserves text performance comparable to text-only models of the same size while expanding omnimodal capability. The article lists benchmark scores including 82.7 on LongAudioSpan for long audio understanding, 63.4 on OmniVideoBench for audio/video collaborative inference, 28.2 on OmniCap-IF for video caption instruction following, and 89.7 on AliMeeting for Chinese multi-speaker transcription, with Qwen claiming it surpassed Gemini 3.8 Flash across multiple tests.

AIHubMix

Coverage

Aibase details Qwen3.8-Omni-Flash's audio strengths, noting it beat Gemini 3.8 Flash on WildClawBench-MM (71.0 vs 58.9), DailyOmni (85.1 vs 84.0), SpotSoundBench (67.2 vs 39.7), and MMAU (81.8 vs 76.9). Multi-speaker meeting recognition improved dramatically, with AliMeeting speaker error rate dropping from 88.11 to 3.35 and cpWER from 89.61 to 17.18. The piece notes Gemini retains advantages on AgenticVBench (45.0 vs 36.8) and OmniGAIA (78.6 vs 74.0), so parity claims mainly apply to overall audio and audio-video capability. Alibaba Cloud international pricing is listed at $0.15 per million input tokens, $0.016 per million cached tokens, and $0.47 per million output tokens.

Videos about Qwen3.8 Omni Flash

More models around Qwen3.8 Omni Flash