Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
AIHubMix logo

Model details

Qwen3.8 Flash

Qwen3.8 Flash sits inside Alibaba's Qwen model family, a line that has been steadily introducing architectural upgrades meant to balance capability with efficient inference. Recent Qwen releases have showcased mixture-of-experts designs and multimodal support, and the broader family has been used to preview upcoming architectures before full successors arrive, signaling a deliberate rollout strategy aimed at giving developers early access to new techniques.

As a Flash-tier variant in the Qwen3.8 lineup, the model is shaped for responsive, general-purpose deployment, where the family's emphasis on multimodal understanding and tool-oriented reasoning tends to translate well to assistants, content workflows, and developer-facing applications that benefit from quick turnaround on text, image, and video inputs. Its placement in the Qwen3.8 generation also makes it a practical choice for teams that want to align with the trajectory leading toward Qwen4, while still relying on a stable, production-oriented model.

AIHubMixqwen3.8-flashqwen

Quick Info

Powered by
Provider
AIHubMix
Model key
qwen3.8-flash
Release date
Aug 26, 2026
Last updated
Aug 26, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.1126
Output token cost
$0.380025

Limits

Output tokens
131,072 tokens
Context window
1,000,000 tokens

Transparent token rates

Compare Qwen3.8 Flash pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3.8 Flash

AIHubMix

CoverageRelease Notes

Alibaba released Qwen3.8-Flash on August 26, 2026, making it downloadable and publishing its weights. The 125B-parameter model uses an architecture designed for the next-generation Qwen 4 series, significantly reducing training and inference costs. Alibaba positions it as a lower-priced option to drive global adoption of its AI lineup. The release is competitive with Anthropic's Opus 4.6 and DeepSeek's V4-Flash, according to Alibaba. The company has pledged more than 380 billion yuan (about US$57 billion) over three years toward AI development and infrastructure, funding the effort in part with a US$10.2 billion Hong Kong follow-on share offering announced the same week.

AIHubMix

CoveragePreview

Qwen3.8-Flash is a 125B-parameter multimodal MoE with 51B n-gram embeddings and 6B active parameters per token, trained at roughly one-ninth the cost of Qwen3.7-Plus. It ships as an early look at the Qwen4 architecture, featuring GDN+QSA hybrid attention, Gated Residual, n-gram embeddings, and Muon training, with QSA delivering up to 7.6x faster 1M-token prefill. The model scores 58.7 on DeepSWE 1.1, 62.5 on SWE-bench Pro, and 73.9 on CoWorkBench, beating Qwen3.7-Plus on coding and office work benchmarks. Weights are available on Hugging Face and ModelScope with day-0 vLLM and SGLang support, and QwenCloud API pricing is $0.16 per million input tokens and $0.47 per million output tokens.

AIHubMix

CoverageBenchmark

The LLM-Stats scorecard for Qwen3.8 Flash, dated August 26, 2026, places the model at $0.17 per million blended tokens with an LLM Stats Score of 49.1, positioning it between DeepSeek-V4-Flash-0731 ($0.066, score 44.7) and GPT-6 Astra ($11.9, score 59.9) on the cost-efficiency chart. Quality Tracker signals show the mo Performance-by-conversation-depth visualization on the same page tracks how Qwen3.8 Flash holds up as conversations get longer, with the chart spanning September 1 through September 14, 2026, ranging between -6.0σ and +6.0σ and showing a positive trajectory over the period. The scorecard explicitly identifies the subje

AIHubMix

CoverageRelease Notes

Alibaba released Qwen3.8-Flash on August 26, 2026, as an open-weight, multimodal Mixture-of-Experts model with 125B parameters plus 51B N-gram embeddings and 6B parameters activated per token. The model serves as an early architecture preview of the upcoming Qwen4 series, targeting high-volume applications, tool-driven workflows, and coding or co-working assistants. Qwen3.8-Flash natively supports 262K tokens of context extensible to 1M, performs competitively against DeepSeek-V4-Flash and Claude-Opus-4.6 on benchmarks including SWE-bench Pro, CoWorkBench, Toolathlon Verified, MathVision, AndroidWorld, and ERQA, and ships weights on Hugging Face and ModelScope with API access at 1 RMB input / 3 RMB output per 1M tokens on Model Studio and Qwen Cloud.

Videos about Qwen3.8 Flash

More models around Qwen3.8 Flash