Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
OpenRouter logo

Model details

Qwen3.8 Omni Flash

The model overview is temporarily unavailable.

OpenRouterqwen/qwen3.8-omni-flashqwen

Quick Info

Powered by
Provider
OpenRouter
Model key
qwen/qwen3.8-omni-flash
Release date
Sep 17, 2026
Last updated
Sep 17, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.15
Output token cost
$0.47

Limits

Output tokens
131,072 tokens
Context window
1,000,000 tokens

Transparent token rates

Compare Qwen3.8 Omni Flash pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3.8 Omni Flash

CrossModel

CoverageBenchmark

The Decoder reports that Qwen3.8-Omni-Flash is Qwen's first multimodal model built for AI agents, processing audio and video together to draw conclusions and edit vlogs, translate short videos, or summarize movies on its own. The context window spans one million tokens, and Qwen positions the model as coming close to m The article describes an agentic video-editing demo in which Qwen3.8-Omni-Flash watches a vlog, plans an edit, and uses tools to deliver a finished short, alongside the open-source Qwen-MM-Plugins that add video editing, speaker recognition, and PDF video notes to agents like Claude Code, Gemini CLI, and Qwen Code, wit

CrossModel

Coverage

Qwen3.8-Omni-Flash targets workflows including video editing, music video creation, film commentary, audio-visual summarization, and real-time conversations, with a 1M-token context window. Qwen reports an average score improvement exceeding 25% across 29 evaluations versus Qwen3.5-Omni-Plus, and claims audio input API costs fell over 98% while audio-visual input costs dropped over 93%. For long videos, the model's agent selectively gathers relevant evidence rather than processing every frame, boosting OmniVideoBench accuracy from 63.4 to 67.8 while cutting token use ~45.7%. It also supports up to one hour of audio-visual meeting input, identifying speakers, transcribing, generating minutes, extracting action items, and triggering downstream tool actions such as email or coding tasks based on meeting context.

CrossModel

CoverageBenchmark

BenchLM's profile for Qwen3.8-Omni-Flash compiles an audio and video benchmark table including WildClawBench-MM 71.0, UniClawBench 69.6, OmniVideoBench 63.4 (static) and 67.8 (Qwen Code), StreamingBench 80.8, AliMeeting DER 3.4 / cpWER 17.2, and VoiceBench 91.6. The page notes the model is ranked 29 on Instruction Foll Decision snapshots on the same page report a 1M-token context window, an unranked capability score (field median 56.4), and explicitly mark the first-party API token rate as not published and time-to-first-token as not measured. Coverage is categorized as 2/2 verified in agentic benchmarks, 5/5 in coding, 8/8 in multim

CrossModel

CoverageBenchmark

DataCamp's overview frames Qwen3.8-Omni-Flash as Alibaba's next-generation omnimodal replacement for Qwen3.5-Omni-Plus, sitting in the Flash tier as a cost-efficient, high-throughput option. It accepts text, image, audio, and video within a 1M-token context, claims audio-visual performance close to Gemini 3.8 Flash and The article highlights headline agentic gains of 36.5 points on WildClawBench-MM and a 98%+ drop in the per-hour price of audio input, and recommends the model for audio- or video-heavy workloads while noting that Qwen3.8-Flash-Next scores marginally higher on pure text or code tasks. It is positioned as undercutting m

CrossModel

CoverageRelease Notes

MarkTechPost covers the September 18, 2026 release of Qwen3.8-Omni-Flash, described as Qwen's first omni-modal model built around agentic capabilities that accepts text, images, audio, and video and returns text. The article confirms it is built on the Qwen3.8-Flash-Next architecture (which shipped with open weights in The same coverage states audio-video understanding, reasoning, and tool use sit inside one model following a "understand, plan, execute, deliver" workflow. The API supports both DashScope and OpenAI protocols, working with Chat Completions and the Responses API, and the model is live on QwenCloud, Alibaba Cloud Model S

CrossModel

CoverageBenchmark

Alibaba's Qwen team released Qwen3.8-Omni-Flash on September 18, 2026, positioning it as the first omnimodal model built for agentic audio-video workflows. According to the explainx.ai launch recap, it accepts text, images, audio, and video within a 1-million-token context window and is framed as shifting from describi The same coverage details a 25%+ average improvement over Qwen3.5-Omni-Plus across 29 evaluations, audio input API pricing reduced more than 98% versus its predecessor, and 45.7% fewer tokens per query on long-video tasks. It also ships two open-source companions — Qwen-MM-Plugins for agent harnesses like Claude Code,

CrossModel

Coverage

Qwen launched Qwen3.8-Omni-Flash on September 18, 2026 as a native multimodal model supporting text, image, audio, and video input within a 1M-token context window. Available on the Qwen AI platform, it expands agentic capabilities into audio- and video-centric workflows including video editing, MV creation, film production, and audio/video summarization. Across 30 evaluations, Qwen reports an average improvement of over 26% versus Qwen3.5-Omni-Plus, with notable gains on WildClawBench-MM (+36.5), AgenticVBench (+22.3), and UniClawBench (69.6). In Agentic Understanding mode, OmniVideoBench accuracy rose from 63.4 to 67.8 while token consumption dropped ~45.7%, and the team open-sourced Qwen-Live Harness to support long workflows and real-time interaction.

Videos about Qwen3.8 Omni Flash

More models around Qwen3.8 Omni Flash