Sulat.com
AI models
NaN logo

Model details

Qwen3.8 Flash

Qwen3.8 Flash is the official production variant built on top of Qwen3.8-Flash-Next, an experimental preview of the architecture expected to underpin the next major Qwen generation. The underlying Next checkpoint is a multimodal Mixture-of-Experts model with around 125B total parameters and roughly 6B parameters activated per token, paired with a 51B n-gram memory module and a small multi-token prediction head used for speculative decoding. Its hybrid attention design replaces the earlier Gated DeltaNet plus gated attention pairing with Gated DeltaNet combined with Qwen Sparse Attention, which selects whole micro-blocks rather than individual tokens to cut long-context latency. Qwen3.8 Flash then layers production conveniences on top of that experimental base, shipping with a 1M-token default context window, official built-in tools, and managed inference through Qwen Cloud.

In practical terms, Qwen3.8 Flash is aimed at teams that want the architectural innovations of the Next preview in a more deployment-ready form. The Next weights are released in Hugging Face Transformers format and are compatible with vLLM, SGLang, and TokenSpeed, so the same model can be self-hosted or consumed through the hosted Qwen Cloud API. The combination of sparse activation, n-gram memory, and micro-block sparse attention is intended to make very long context affordable, while the multimodal inputs and tool-calling capabilities make it suitable for assistants, document analysis, and structured-output workflows that benefit from large effective context without the cost of a fully dense model.

NaNqwen3.8-flashqwen

Quick Info

Powered by
Provider
NaN
Model key
qwen3.8-flash
Release date
Aug 26, 2026
Last updated
Aug 26, 2026
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
131,072 tokens
Context window
262,144 tokens

Latest news about Qwen3.8 Flash

NaN

CoverageBenchmark

Alibaba's Qwen team released Qwen3.8-Flash-Next on August 26, 2026, an open-weight 125B multimodal mixture-of-experts model with 6B active parameters per token that serves as an early architectural preview of the upcoming Qwen4 family, much as Qwen3-Next did for Qwen3.5. The DataCamp explainer details native 262,144-to Benchmark results published in the article show Qwen3.8-Flash-Next scoring 62.5 on SWE-bench Pro (versus Claude Opus 4.6 Max's 53.4), 73.9 on CoWorkBench (versus 68.2), and 55.7 on JobBench (versus 36.6), indicating competitive coding and agent performance against premium proprietary models. On Humanity's Last Exam it

Videos about Qwen3.8 Flash

More models around Qwen3.8 Flash