Currently listed through these providers:
Model details
Qwen 3.8 Flash
Qwen 3.8 Flash Next is part of Alibaba's evolving Qwen family and is described in third-party coverage as an open-weight preview pointing toward the upcoming Qwen4 generation. The model is listed on ModelScope under the Qwen namespace with a 125B total / 6B active parameter profile, indicating a Mixture-of-Experts design where only a fraction of the parameters are engaged per token. This sparse activation pattern is the architectural choice that gives the "Next" variant its name, aiming to deliver stronger reasoning capacity while keeping inference costs closer to those of smaller dense models.
As a Flash-class release, the model is positioned for practical agent and assistant workloads rather than pure research benchmarks, with community discussion highlighting its appeal for local high-memory setups. Coverage frames it as Alibaba's open-weight answer in the mid-to-large reasoning tier, building on the lineage of the Qwen3 series and signaling the direction the Qwen4 generation will take. Developers evaluating the model can expect a familiar Qwen-style chat experience with broader capability, suitable for tasks that benefit from larger context handling and tool-oriented workflows.
Quick Info
Powered by- Provider
- Venice AI
- Model key
- qwen-3-8-flash
- Release date
- Sep 10, 2026
- Last updated
- Sep 10, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.14
- Output token cost
- $0.49
Limits
- Output tokens
- 131,072 tokens
- Context window
- 1,000,000 tokens
Transparent token rates
Compare Qwen 3.8 Flash pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about Qwen 3.8 Flash
No articles yet. Fetch the latest news to show it here.