Currently listed through these providers:
Model details
Qwen3.8 Flash
Qwen3.8 Flash sits inside the broader Qwen3 family as a multimodal Mixture-of-Experts release from Alibaba, framed around an "innovative model architecture" and "optimal price-performance" in the official Alibaba Cloud Community announcement. The model is designed for practical, cost-sensitive deployments that still require reasoning and tool-use capabilities, which is reflected in the deployment ecosystem that has formed around it. A closely related artifact, Qwen3.8-Flash-Next, is already published on Hugging Face by Inferact with quantized weights in NVFP4, FP8, and BF16 formats, indicating an active community focus on efficient inference variants of the Flash line.
From a practical fit standpoint, the vLLM recipe for the related Flash variant demonstrates how the model is meant to be served: it runs on a wide range of NVIDIA and AMD accelerators including H100, H200, B200, GB200, GB300, RTX Pro 6000, and AMD MI300X, MI325X, and MI355X hardware. The recipe configures the Qwen3-specific tool-calling and reasoning parsers, and exposes configuration knobs for prefix caching, batched tokens, and GPU memory utilization, suggesting the model is engineered for latency-aware serving with flexible throughput tuning. These signals point to Qwen3.8 Flash being best suited for teams that need a multimodal, reasoning-capable model with efficient serving characteristics and the ability to plug into agent-style tool workflows.
Quick Info
Powered by- Provider
- OpenRouter
- Model key
- qwen/qwen3.8-flash
- Release date
- Aug 26, 2026
- Last updated
- Aug 26, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.15
- Output token cost
- $0.47
Limits
- Output tokens
- 131,072 tokens
- Context window
- 1,000,000 tokens
Latest news about Qwen3.8 Flash
Videos about Qwen3.8 Flash
Recent tweets and retweets from OpenRouter
More models around Qwen3.8 Flash
This exact model name is also listed by 14 other providers.