Sulat.com
AI models
Requesty logo

Model details

Qwen3.8 Flash

Qwen3.8 Flash is positioned as a production-ready build that sits on top of the Qwen3.8-Flash-Next experimental release, adding the long context window and managed tooling that distinguish a hosted endpoint from a research artifact. The model uses a mixture-of-experts design with roughly 125B total parameters and activates only a small fraction per token, paired with an unusually sized 51B n-gram embedding component that acts as a secondary memory system alongside a small multi-token prediction head used for speculative decoding. Alibaba openly framed the architecture as an early look at the design that will carry forward into Qwen4, marking it as more of a directional preview than a conventional incremental release within the existing lineup.

In practical terms, the hosted variant of Qwen3.8 Flash is aimed at workloads that need to chew through very long inputs without sacrificing multimodal coverage, since the production configuration defaults to a one-million-token context while still accepting text, image, and video inputs. Independent reporting noted that the underlying Next variant was shown running locally on about 75GB of RAM with day-zero Unsloth support, suggesting that the same architecture can be self-hosted on a single high-memory workstation for developers who prefer to keep inference in-house. Combined with budget-tier token pricing and tool calling support, the model is a reasonable fit for long-context document analysis, multimodal assistants, and agent pipelines where an open-weights lineage matters but a managed, production-grade endpoint is preferred.

Requestyqwen3.8-flashqwen

Quick Info

Powered by
Provider
Requesty
Model key
qwen3.8-flash
Release date
Aug 26, 2026
Last updated
Aug 26, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.16
Output token cost
$0.47

Limits

Output tokens
131,072 tokens
Context window
1,048,576 tokens

Transparent token rates

Compare Qwen3.8 Flash pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3.8 Flash

Requesty

Official sourceAnnouncement

Requesty's August 28, 2026 blog post directly covers the Qwen3.8-Flash launch by Alibaba on August 26, 2026, describing it as an open-weight multimodal mixture-of-experts model with 125B parameters plus a 51B N-gram component. It is framed as an early preview of the Qwen4 architecture, with the production API priced at The post adds concrete deployment details: Alibaba demonstrated the 125B model running locally on 75GB of RAM with day-zero Unsloth support, and it was one of five open-weight releases in nine days alongside GLM-5.3-Flash, Hy4 Preview, MiniMax M3/M2.7, and DeepSeek V4-Flash-Vision-Exp. This makes Qwen3.8-Flash part of

Requesty

Official sourceBenchmark

Requesty's live model catalog lists qwen3.8-flash under Alibaba (Qwen) at $0.16 input and $0.47 output per million tokens, with a 1.0M-token context window, currently served through one provider on the Requesty router. It appears alongside related entries including qwen3.8-flash-next (262K context, $0.20/$0.50) and qwe The catalog frames the offering as part of a 700+ model, 33-provider OpenAI-compatible routing layer covering 31 labs, with Flash-class models from Alibaba, Z.AI, DeepSeek, Google, and others priced in the budget tier. It corroborates the blog's pricing figures and confirms live availability of qwen3.8-flash at the doc

Videos about Qwen3.8 Flash

More models around Qwen3.8 Flash