Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Ofox logo

Model details

Qwen3.8 Flash

Qwen3.8 Flash is a multimodal reasoning model from Alibaba positioned for practical assistant workloads rather than pure chat. The OpenRouter listing frames it as well suited to coding assistance, agentic workflows, visual understanding, document and codebase analysis, desktop interaction, chart analysis, and long-video analysis, reflecting a design intent that blends language reasoning with perception over text, images, and video. That broad intent makes it a reasonable choice for builders who want a single model behind pipelines that mix tool use, structured outputs, and on-screen content.

On the architecture side, the model sits in the Qwen3 family as a redesigned multimodal mixture-of-experts design with 125B main parameters, an extra 51B of N-gram embeddings, and about 6B parameters activated per token, which is what enables efficient training and inference compared with denser configurations. The same announcement highlights comprehensive upgrades across attention, residual connections, embeddings, and optimization, pointing to architectural innovations intended to push capability boundaries while keeping compute costs manageable. For practitioners, the combination of MoE sparsity, multimodal coverage, and a very long context window suggests a model meant for long documents and video analysis where efficiency and breadth of understanding matter more than peak single-token reasoning depth.

Ofoxbailian/qwen3.8-flashqwen

Quick Info

Powered by
Provider
Ofox
Model key
bailian/qwen3.8-flash
Release date
Aug 26, 2026
Last updated
Aug 26, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.15
Output token cost
$0.47

Limits

Output tokens
131,072 tokens
Context window
1,000,000 tokens

Transparent token rates

Compare Qwen3.8 Flash pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3.8 Flash

Ofox

Coverage

The QCode status tracker documents Qwen3.8-Flash as released on 2026-08-26, with a hosted list price on QwenCloud of $0.16 per million input tokens and $0.47 per million output tokens and an advertised 1-million-token context window. The page distinguishes the production API SKU (Qwen3.8-Flash) from the architecture-pr An Alibaba Cloud Bailian notice dated 2026-08-26 23:18 set adjusted RMB pricing effective 2026-08-27 12:00 Beijing time, lowering Qwen3.8-Flash input from ¥1.00 to ¥0.80 per million tokens and output from ¥3.00 to ¥2.70 per million tokens, while international pages continue to quote $0.16/$0.47 in USD. The tracker plac

Ofox

CoverageBenchmark

DataCamp published a detailed explainer on August 27, 2026, about Qwen3.8-Flash-Next, the open-weight 125B mixture-of-experts multimodal model that Alibaba's Qwen team released on August 26, 2026 as an early architectural preview of the upcoming Qwen4 family. The article emphasizes cost efficiency, noting the new archi The explainer situates Qwen3.8-Flash-Next as a cost-efficient workhorse below Max-class flagships, explicitly framing it as a "Flash" tier release whose architectural changes preview Qwen4 ahead of the full model lineup. Readers are pointed to companion tutorials for running the model locally and for context on related

Ofox

Coverage

Alibaba's Qwen team released Qwen3.8-Flash on August 26, 2026, positioning it as a new multimodal model with stronger coding and office-task performance while significantly reducing training costs. The model features a default context window of 262,144 tokens that is expandable to 1 million tokens, enabling it to handl Qwen announced API pricing of 1 yuan (approximately $0.1488) per million input tokens and 3 yuan per million output tokens for access through its application programming interface. Alongside the production release, Qwen also open-sourced weights for Qwen3.8-Flash-Next, which Alibaba describes as an architecture prototy

Ofox

CoverageBenchmark

AI Release Tracker logged Qwen3.8-Flash-Next as an open-weight model released by Qwen on Wednesday, August 26, 2026, twelve days after Qwen3.8-27B. The entry specifies a 125B-parameter architecture with a 262K-token context window and aggregates benchmark results across IFBench, Agent's Last Exam, SWE-Bench Pro, SWE-Be Agentic and tool-use benchmarks include Agent's Last Exam at 24.3% pass@1 with a 51.2% score (6 of 7, best GPT-6 Astra at 59.3% score) and JobBench at 55.7% (3 of 8, best Muse Spark 1.3 at 64.9%), with CoWorkBench long-horizon office-work results also tracked. The tracker does not address the Ofox gateway, pricing, or

Videos about Qwen3.8 Flash

More models around Qwen3.8 Flash