Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
TensorX logo

Model details

Qwen3.8 Flash Next

Qwen3.8 Flash Next is positioned as an early architectural preview of the upcoming Qwen4 family, following the same pattern the Qwen team used when Qwen3-Next previewed changes ahead of Qwen3.5. The model is released with open weights and is described as a multimodal mixture-of-experts design intended to deliver strong general capability at substantially lower training cost than its predecessor, Qwen3.7-Plus. This combination of openness and cost efficiency targets developers who want to experiment with next-generation Qwen behavior locally before the full Qwen4 line arrives.

In benchmark performance, Qwen3.8 Flash Next ranks competitively across several tracked categories, placing in the top ten percent for tool calling, within the top tier for vision and reasoning, and in the upper half for coding and long-context tasks. It scores 0.96 on MathVision and is reported to outperform heavyweight models such as Claude Opus 4.6 Max on SWE-bench coding evaluations, suggesting it is well suited for coding assistants and multimodal reasoning workflows where both image understanding and tool use matter. Practically, the model fits users who need an affordable, open-weights entry point to the forthcoming Qwen4 generation and are comfortable with a preview-stage release.

TensorXqwen/qwen3.8-flash-nextqwen

Quick Info

Powered by
Provider
TensorX
Model key
qwen/qwen3.8-flash-next
Release date
Aug 27, 2026
Last updated
Aug 27, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.20
Output token cost
$0.50

Limits

Output tokens
64,000 tokens
Context window
262,144 tokens

Transparent token rates

Compare Qwen3.8 Flash Next pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3.8 Flash Next

TensorX

Coverage

An NVIDIA developer-forum thread posted September 8, 2026 by community member azampatti documents an INT4 AutoRound quantization recipe for Qwen3.8-Flash-Next targeting DGX Spark / GB10 hardware. The community fork halves the number of routed experts per token from 10 to 5 and heals the resulting quality loss with a 37 The reported throughput is 60–70 tokens per second in coding workloads, with vLLM launching a GPU KV cache of approximately 644,732 tokens and a maximum concurrency of about 2.46× at 262,144 tokens per request. The variant name Qwen3.8-Flash-Next-125B-A5B-INT4-AutoRound indicates this is a community-modified derivative

TensorX

CoverageBenchmark

A Kaitchup (Substack) review dated August 27, 2026 describes Qwen3.8-Flash-Next as a multimodal Mixture-of-Experts model with a 125B-parameter main model, 6B activated per token, a 51B n-gram embedding system, and a 4B MTP module for speculative decoding. It notes the Transformers configuration already identifies the a The review plans to compare Qwen3.8-Flash-Next with Qwen3.8 27B, dig into the Qwen Sparse Attention, Gated Residual, and 51B-parameter n-gram embedding table, and estimate memory consumption at 256K context including NVFP4 and Unsloth GGUF quantizations. Substantive sections are gated behind the paid tier, so only the

TensorX

CoveragePreview

An Enera enterprise-focused analysis dated August 27, 2026 frames Qwen3.8-Flash-Next as a 176-billion-parameter multimodal MoE model and official preview of the Qwen4 architecture, following the same playbook Qwen used with Qwen3-Next before the Qwen3.5 series. It highlights that the activated-parameter figure of 6B dr The post details four systematic upgrades across attention (GDN paired with Qwen Sparse Attention, with three out of four layers using GDN for history compression), residual connections (Gated Residual modulating widened residual streams), embedding (the 51B n-gram table as a cheaper-than-MoE scaling axis), and the tra

TensorX

CoveragePreview

A Local AI Zone technical deep-dive dated August 27, 2026 characterizes Qwen3.8-Flash-Next as Alibaba's second-generation "Next" architecture preview and the open-weight counterpart to the production qwen3.8-flash API SKU. It confirms the Transformers config identifies the architecture as qwen4_exp and walks through th The deep-dive reports self-reported headline scores including DeepSWE 1.1 at 58.7 (vs Qwen3.7-Plus 16.5), SWE-bench Pro at 62.5, CoWorkBench at 73.9, GPQA Diamond at 91.7, LiveCodeBench v6 at 91.9, and AndroidWorld at 84.5, while acknowledging losses on NL2Repo-Bench, HLE, and OSWorld 2.0 binary. It cites QwenCloud (Da

TensorX

CoverageBenchmark

The llm-stats model page for Qwen3.8-Flash-Next aggregates benchmark performance and capability-tier rankings, listing scores sourced from the model's scorecard, paper, or official blog posts. It places the model at rank 20 on the LLM Stats Score composite, with capability tiers of Top 10% in Tool Calling (11 of 194), Specific benchmark rankings cited include MathVision at rank 3 with 0.96 (with code interpreter methodology), LiveCodeBench v6 at rank 1 with 0.92, GPQA Diamond at rank 20 with 0.92, and CharXiv-R at rank 4 with 0.91, with scores sourced from huggingface.co. The page serves as a consolidated leaderboard view of the mod

TensorX

Coverage

An llm-stats launch summary dated August 26, 2026 catalogs Qwen3.8-Flash-Next under id qwen3.8-flash-next with Hugging Face repo Qwen/Qwen3.8-Flash-Next, released under the qwen-community-1.0 license on the HF card. The post reiterates the official positioning: this is an open-weight Qwen4-architecture preview, not the Architecture figures list 180B stored parameters broken into a 6B-active main model plus a 51B N-gram table plus 4B MTP, with 262,144 native context extensible to ~1M via YaRN. Self-reported deltas versus DeepSeek-V4-Flash-0731 show Qwen3.8-Flash-Next ahead on DeepSWE 1.1 (58.7 vs 54.4), SWE-bench Pro (62.5 vs 56.0), T

TensorX

Coverage

Alibaba's Qwen team open-sourced Qwen/Qwen3.8-Flash-Next on August 26, 2026 as an experimental preview of the architecture that will underpin Qwen4, distributed in Hugging Face Transformers format and compatible with vLLM, SGLang, and TokenSpeed. The model card explicitly distinguishes this open-weight release from the Architecturally, the model introduces three headline innovations: Qwen Sparse Attention (QSA), which pairs Gated DeltaNet with a micro-block-level sparse attention that targets long-context latency for agentic workloads; Gated Residual, which adds a data-dependent read gate and per-branch scalar write gate around widen

Videos about Qwen3.8 Flash Next

More models around Qwen3.8 Flash Next