Currently listed through these providers:
Model details
Qwen3.8 Flash Next
Qwen3.8 Flash Next is positioned by its developers as an experimental preview of the architecture intended to underpin Qwen4, making it a useful early look at the family's next generation rather than a fully polished production release. The model combines a 125B-parameter Mixture-of-Experts backbone with only 6B parameters activated per token, a 51B-parameter n-gram embedding system for memory, and a 4B Multi-Token Prediction module used for speculative decoding, illustrating a hybrid approach that blends sparse expert routing with auxiliary retrieval-style components. Because the configuration files already identify the architecture as qwen4_exp, the release signals Qwen's direction toward more efficient scaling rather than simply larger dense models, and the open weights make that architectural shift directly inspectable for researchers and local inference practitioners.
In practical terms, the model targets users who want to experiment with cutting-edge hybrid attention designs, particularly the Gated DeltaNet paired with Qwen Sparse Attention, which operates at the micro-block level to reduce long-context latency. It accepts both text and image input while producing text output, and the weights are published on Hugging Face under the Qwen organization with compatibility for Transformers, vLLM, SGLang, and TokenSpeed, so local and self-hosted workflows are a natural fit. A separate managed Qwen3.8-Flash production version with a longer 1M-token context and built-in tools is offered through Qwen Cloud for teams that prefer hosted inference, while this Flash Next variant is better suited to architecture exploration, benchmarking, and prototyping around the Qwen4 design choices.
Quick Info
Powered by- Provider
- Cortecs
- Model key
- qwen3.8-flash-next
- Release date
- Aug 27, 2026
- Last updated
- Aug 27, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.201
- Output token cost
- $0.50
Limits
- Output tokens
- 64,000 tokens
- Context window
- 262,144 tokens