Currently listed through these providers:
Model details
Qwen3.8 Flash
Within the Qwen family of models, Qwen3.8 Flash continues a line of fast, multimodal offerings designed for practical text-and-image interactions at scale. It is positioned as a general-purpose assistant that can take in text and image inputs and produce text output, with explicit support for reasoning, tool calling, structured output, and working with attachments. That combination makes it suitable for workflows that need more than a plain chat reply, such as extracting fields from documents, driving API calls through tools, and returning results in a defined JSON shape for downstream software.
Coverage from third-party write-ups of the closely related Qwen3.8-Flash-Next build points to a multimodal mixture-of-experts design with a 125B-parameter main model and only 6B parameters active per token, paired with a 51B n-gram embedding system and a 4B speculative-decoding module, presented as an architectural preview of the upcoming Qwen4 generation. Reviews highlight aggressive efficiency gains — reported as roughly a ninth of the training cost of its predecessor — alongside strong coding results on SWE-bench Pro, where it is described as outperforming Claude Opus 4.6 Max. For practitioners, the practical takeaway is a flash-tier model aimed at routine reasoning and coding assistance at a low per-token price, with a very large context window that suits long-document analysis and multi-step agentic tasks.
Quick Info
Powered by- Provider
- Charm Hyper
- Model key
- qwen3.8-flash
- Release date
- Aug 26, 2026
- Last updated
- Aug 26, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.15
- Output token cost
- $0.47
Limits
- Output tokens
- 128,000 tokens
- Context window
- 1,000,000 tokens
Latest news about Qwen3.8 Flash
Videos about Qwen3.8 Flash
More models around Qwen3.8 Flash
This exact model name is also listed by 14 other providers.