Currently listed through these providers:
Model details
Qwen3.8 Flash
Built around a multimodal mixture-of-experts design, this model was released by Alibaba's Qwen team as an early architectural preview of the upcoming Qwen4 family, following the same preview-then-productize pattern that Qwen3-Next played for Qwen3.5. The core innovation is a hybrid attention stack pairing Gated DeltaNet (GDN) compression with Qwen Sparse Attention (QSA), supported by systematic upgrades to residual connections, embedding, and optimization that together push capability and training stability forward while trimming compute. Because it is an MoE model, only a slice of the total parameters activates per token, which is the lever behind the dramatic drop in training cost relative to the prior Qwen3.7-Plus generation.
In practice the model behaves like a fast, code-leaning flagship: it tops major software engineering evaluations such as SWE-bench Pro and posts the leading score on CoWorkBench, outperforming heavier closed competitors on those coding workloads. The combination of a hybrid GDN + QSA attention backbone and aggressive MoE sparsity makes it a strong fit for production assistants that need long-context reasoning, agentic tool use, and structured outputs at a fraction of the training and inference cost of dense predecessors. For teams planning around the Qwen4 line, it is best understood as a forward-looking testbed: the architectural gains seen here are the same gains that will define the next generation of Qwen deployments.
Quick Info
Powered by- Provider
- Vancine
- Model key
- qwen3.8-flash
- Release date
- Aug 26, 2026
- Last updated
- Aug 26, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.12
- Output token cost
- $0.38
Limits
- Output tokens
- 131,072 tokens
- Context window
- 1,000,000 tokens
Latest news about Qwen3.8 Flash
Videos about Qwen3.8 Flash
More models around Qwen3.8 Flash
This exact model name is also listed by 14 other providers.