Sulat.com
AI models
OpenRouter logo

Model details

Qwen3.8 Flash

Qwen3.8 Flash sits inside the broader Qwen3 family as a multimodal Mixture-of-Experts release from Alibaba, framed around an "innovative model architecture" and "optimal price-performance" in the official Alibaba Cloud Community announcement. The model is designed for practical, cost-sensitive deployments that still require reasoning and tool-use capabilities, which is reflected in the deployment ecosystem that has formed around it. A closely related artifact, Qwen3.8-Flash-Next, is already published on Hugging Face by Inferact with quantized weights in NVFP4, FP8, and BF16 formats, indicating an active community focus on efficient inference variants of the Flash line.

From a practical fit standpoint, the vLLM recipe for the related Flash variant demonstrates how the model is meant to be served: it runs on a wide range of NVIDIA and AMD accelerators including H100, H200, B200, GB200, GB300, RTX Pro 6000, and AMD MI300X, MI325X, and MI355X hardware. The recipe configures the Qwen3-specific tool-calling and reasoning parsers, and exposes configuration knobs for prefix caching, batched tokens, and GPU memory utilization, suggesting the model is engineered for latency-aware serving with flexible throughput tuning. These signals point to Qwen3.8 Flash being best suited for teams that need a multimodal, reasoning-capable model with efficient serving characteristics and the ability to plug into agent-style tool workflows.

OpenRouterqwen/qwen3.8-flashqwen

Quick Info

Powered by
Provider
OpenRouter
Model key
qwen/qwen3.8-flash
Release date
Aug 26, 2026
Last updated
Aug 26, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.15
Output token cost
$0.47

Limits

Output tokens
131,072 tokens
Context window
1,000,000 tokens

Latest news about Qwen3.8 Flash

Videos about Qwen3.8 Flash

Recent tweets and retweets from OpenRouter

More models around Qwen3.8 Flash