Currently listed through these providers:
Model details
Qwen3.8 Flash
This model is positioned as a multimodal mixture-of-experts design intended as an early architectural preview of the next major Qwen generation, following the same playbook that earlier "Next" releases used to introduce structural changes ahead of a full model line. The published description frames it as a 125B-parameter MoE with roughly 6B active parameters per token, an efficient routing scheme that reportedly brought training costs down to about a ninth of its predecessor while preserving competitive capability. Its multimodal input combined with text-only output makes it suitable for workflows that need to ingest images and video frames alongside text, then produce structured or conversational text in return.
In practical terms, the model is aimed at coding and reasoning-heavy applications that also benefit from large context windows. Reported head-to-head results show it outperforming established frontier coding models on benchmarks such as SWE-bench Pro (62.5 versus 53.4) and CoWorkBench (73 versus a leading competitor), suggesting a particular fit for software engineering agents, tool-using assistants, and long-document analysis. Deployment guidance emphasizes single-node tensor parallelism across modern accelerators, with NVFP4 quantized checkpoints available for teams that want to trade a small amount of precision for memory efficiency, making the model approachable for production inference on current-generation GPU clusters.
Quick Info
Powered by- Provider
- CrossModel
- Model key
- qwen/qwen3.8-flash
- Release date
- Aug 26, 2026
- Last updated
- Aug 26, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.13
- Output token cost
- $0.43
Limits
- Output tokens
- 131,072 tokens
- Context window
- 1,000,000 tokens
Latest news about Qwen3.8 Flash
Videos about Qwen3.8 Flash
More models around Qwen3.8 Flash
This exact model name is also listed by 14 other providers.