Sulat.com
AI models
Cortecs logo

Model details

Qwen3.8 Flash Next

Qwen3.8 Flash Next is positioned by its developers as an experimental preview of the architecture intended to underpin Qwen4, making it a useful early look at the family's next generation rather than a fully polished production release. The model combines a 125B-parameter Mixture-of-Experts backbone with only 6B parameters activated per token, a 51B-parameter n-gram embedding system for memory, and a 4B Multi-Token Prediction module used for speculative decoding, illustrating a hybrid approach that blends sparse expert routing with auxiliary retrieval-style components. Because the configuration files already identify the architecture as qwen4_exp, the release signals Qwen's direction toward more efficient scaling rather than simply larger dense models, and the open weights make that architectural shift directly inspectable for researchers and local inference practitioners.

In practical terms, the model targets users who want to experiment with cutting-edge hybrid attention designs, particularly the Gated DeltaNet paired with Qwen Sparse Attention, which operates at the micro-block level to reduce long-context latency. It accepts both text and image input while producing text output, and the weights are published on Hugging Face under the Qwen organization with compatibility for Transformers, vLLM, SGLang, and TokenSpeed, so local and self-hosted workflows are a natural fit. A separate managed Qwen3.8-Flash production version with a longer 1M-token context and built-in tools is offered through Qwen Cloud for teams that prefer hosted inference, while this Flash Next variant is better suited to architecture exploration, benchmarking, and prototyping around the Qwen4 design choices.

Cortecsqwen3.8-flash-nextqwen

Quick Info

Powered by
Provider
Cortecs
Model key
qwen3.8-flash-next
Release date
Aug 27, 2026
Last updated
Aug 27, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.201
Output token cost
$0.50

Limits

Output tokens
64,000 tokens
Context window
262,144 tokens

Latest news about Qwen3.8 Flash Next

Videos about Qwen3.8 Flash Next

More models around Qwen3.8 Flash Next