Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Vultr logo

Model details

Qwen3.8 Flash Next

Qwen3.8 Flash Next arrives as a release oriented toward developers running large models on consumer and prosumer hardware with substantial unified memory. Community discussion on the NVIDIA developer forums describes the model as explicitly targeting Mac systems in the 96–128GB range, NVIDIA DGX Spark, and AMD Strix Halo 128GB configurations, signaling a design philosophy that prioritizes fitting into workstation-class memory budgets rather than competing at the very high end of parameter counts. The structure of the model has drawn informal comparisons to other recent efficient architectures such as Ling 3.0 and Kimi K3 Flash, suggesting it belongs to a wave of models engineered to balance capability with the practical constraints of local inference.

In benchmark aggregation, the model earns a composite score of 64.49 and holds the 39th position across 637 tracked models and 496 benchmarks as recorded at the end of September 2026, with its strongest published evidence concentrated in multimodal and grounded tasks such as screenshot interpretation, document analysis, and chart reasoning. This profile positions it as a capable generalist with a particular edge in grounded multimodal workflows, appealing to practitioners who need reliable document and visual understanding without resorting to the largest frontier-scale models. For teams evaluating deployment targets, the combination of memory-friendly scaling and solid multimodal grounding makes it a practical fit for workstation-local experimentation and for production paths that can take advantage of FP8 inference on supported hardware.

Vultrqwen3.8-flash-nextqwen

Quick Info

Powered by
Provider
Vultr
Model key
qwen3.8-flash-next
Release date
Aug 27, 2026
Last updated
Aug 27, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.10
Output token cost
$0.20

Limits

Output tokens
131,072 tokens
Context window
262,144 tokens

Transparent token rates

Compare Qwen3.8 Flash Next pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3.8 Flash Next

Vultr

Coverage

Qwen3.8-Flash-Next 176B received day-0 support on NVIDIA NIM, per an NVIDIA Developer Forums announcement dated August 26, 2026. The model is a multimodal mixture-of-experts design with a 125B-parameter main model plus 51B N-gram embeddings, activating 6B parameters per token. It has a native 262,144-token context window, extensible to 1M tokens with YaRN. On GB300 NVL72 (72 Blackwell Ultra GPUs, 130 TB/s NVLink) the FP8 variant with TensorRT-LLM reaches 16,000 tokens/sec per GPU and 200 tokens/sec per user. The model also runs on DGX Spark clusters, DGX Station, and 4x RTX PRO 6000 Blackwell workstations. Day-0 supported stacks include SGLang, vLLM, TensorRT-LLM, NeMo AutoModel SFT + LoRA, and NeMo RL recipes, with weights on Hugging Face and ModelScope.

Vultr

CoverageBenchmark

Alibaba's Qwen team released Qwen3.8-Flash-Next on August 26, 2026, an open-weight 125B MoE model activating 6B parameters per token, positioned as an architectural preview of the upcoming Qwen4 family. Training cost is roughly one-ninth that of its predecessor Qwen3.7-Plus while beating Claude Opus 4.6 Max on SWE-bench Pro (62.5 vs 53.4), CoWorkBench (73.9 vs 68.2), and JobBench (55.7 vs 36.6). The model supports a native 262,144-token context window, extensible to 1M tokens with YaRN, delivering up to 8.6x the prefill throughput of Qwen3.7-Plus at 1M tokens. Pricing on QwenCloud is $0.16 per million input tokens and $0.47 per million output tokens under the name qwen3.8-flash. It trails Claude Opus 4.6 Max on Humanity's Last Exam (35.9 vs 40.0) and shows brittleness on very long agent chains.

Videos about Qwen3.8 Flash Next

More models around Qwen3.8 Flash Next