Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
GMI Cloud logo

Model details

Qwen3.8 Flash

Qwen3.8 Flash was introduced by Alibaba's Qwen team as a multimodal model designed to handle text along with visual inputs while delivering stronger coding and office-task performance. According to the launch announcement, the model was trained at roughly one-ninth the cost of the previous Qwen3.7-Plus, signaling an efficiency-focused step in the Qwen lineup rather than a simple scale increase. Its developer-facing positioning centers on serving real productivity workflows such as software development, document handling, and complex research tasks where balanced reasoning and tool use matter more than raw scale.

A defining practical strength is the model's very long working memory: it ships with a default context window that can be extended to handle large files, lengthy conversations, and extensive research materials, which makes it well suited to agent-style and retrieval-heavy applications. Within the broader release, Qwen3.8-Flash-Next was published as an open-weight sibling previewing the next-generation Qwen4 architecture, giving the developer community an early look at the direction of the family. Together, these releases suggest Qwen is iterating toward more cost-efficient multimodal systems that can still manage demanding, long-context workloads.

GMI CloudQwen/Qwen3.8-Flashqwen

Quick Info

Powered by
Provider
GMI Cloud
Model key
Qwen/Qwen3.8-Flash
Release date
Aug 26, 2026
Last updated
Aug 26, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.16
Output token cost
$0.47

Limits

Output tokens
131,072 tokens
Context window
1,048,575 tokens

Transparent token rates

Compare Qwen3.8 Flash pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3.8 Flash

Ofox

Coverage

This IntuitionLabs technical report covers Qwen3.8-Flash-Next architecture in depth, confirming 125B parameters with 6B activated per token, a 51B n-gram auxiliary table, and a 4B multi-token-prediction module, totaling approximately 180B parameters across 48 layers with native 262,144 context extensible to 1M via YaRN The article addresses deployment memory planning specifically for Qwen3.8-Flash-Next, noting the official FP8 checkpoint is 172.78 GiB (about 185.5 GB) of weight storage rather than full GPU-serving memory, with the vLLM recipe validating FP8 at TP2 minimum on GB300 and recommending TP4, and n-gram embedding requiring

Ofox

CoverageBenchmark

This candidate explicitly names the hosted API model string "qwen/qwen3.8-flash" and ties it to Alibaba's Qwen 3.8 family, describing Qwen 3.8 Flash as a sparse mixture-of-experts multimodal reasoning model with 125B total parameters, 6B active per token, 512 experts, native 262,144-token context extensible to 1M, and The article provides benchmark and capability context specific to Qwen 3.8 Flash, positioning it as the cheap tier of the Qwen 3.8 family suited for product builds, with sibling context distinguishing it from the 27B local-rig variant and the Max variant for budget-unconstrained use. It reports modality support for tex

Ofox

CoverageBenchmark

This buildfastwithai.com review directly covers the Qwen3.8-Flash-Next open-weight checkpoint, which is the foundation of the hosted Qwen3.8 Flash model, reporting that Qwen released weights on August 26, 2026 and that the architecture combines a 125B main model with 51B n-gram embedding parameters, 6B activated per to The article reports Qwen-published benchmark numbers including 91.9 on LiveCodeBench, 62.5 on SWE-bench Pro, 81.0 on SWE-bench Multilingual, 58.7 on DeepSWE 1.1, 73.9 on CoWorkBench, and 55.7 on JobBench, flagging these as provider-reported rather than independently verified. It positions the release as strategically i

Ofox

CoverageBenchmark

This DataCamp article covers Qwen3.8-Flash-Next as an open-weight 125B MoE model with 6B active per token, previewing the Qwen4 architecture, released August 26, 2026, the same day as GLM-5.3 Flash. It explicitly compares the model to Claude Opus 4.6 Max on SWE-bench Pro (62.5 vs 53.4), CoWorkBench (73.9 vs 68.2), and The article frames Qwen3.8-Flash-Next as a cost-efficient workhorse below the Max-class models in raw intelligence but undercutting nearly everything on price, mirroring the Qwen3-Next-to-Qwen3.5 release pattern. Technical details include multimodal support, an experimental Qwen4 architecture preview status, and positi

Ofox

CoverageBenchmark

This llm-stats.com aggregator leaderboard page covers the hosted Qwen3.8 Flash variant explicitly, reporting a composite LLM Stats Score of 48.7 at a blended price of $0.17 per 1M tokens, ranking it 24th overall with cost-efficiency positioning between DeepSeek-V4-Flash-0731 ($0.066) and DeepSeek-V4.1-Flash ($0.24). It The leaderboard data is sourced from the model's Hugging Face card and self-reported provider scorecards, and the page is flagged for captcha verification which can impede direct access, limiting confidence. Pricing appears to be a blended or reseller figure distinct from the QwenCloud direct rate of $0.16 per 1M input

CrossModel

Coverage

Reuters reports that Alibaba's Qwen team released the multimodal Qwen3.8-Flash model on August 26, 2026, positioning it as a cost-efficient upgrade over its predecessor. According to Qwen's statement published on social media, the model supports a default context window of 262,144 tokens, expandable to 1 million tokens The article notes that Qwen would charge 1 yuan (approximately $0.1488) per million input tokens and 3 yuan per million output tokens for access through its application programming interface. In addition, Qwen released open-source weights for a separate Qwen3.8-Flash-Next variant, allowing the developer community to ev

Videos about Qwen3.8 Flash

More models around Qwen3.8 Flash