Sulat.com
AI models
SCNet Token Plan logo

Model details

Qwen3.8 Flash

Qwen3.8 Flash is the production-oriented model in the Qwen family built on the open-weights Qwen3.8-Flash-Next base, which Alibaba describes as an experimental preview of the architecture intended to underpin the next Qwen generation. The base model introduces a redesigned multimodal mixture-of-experts design with 125B main parameters supplemented by 51B of N-gram embeddings and around 6B parameters activated per token, along with architectural upgrades across attention, residual, embedding, and optimization components. Qwen3.8 Flash layers production features on top of this base, including a default 1,000,000-token context window and official built-in tools, making it a natural fit for long documents, full codebases, and extended video or chart analysis where sustained reasoning over very long inputs matters.

In practice, the model is positioned for coding assistance, tool use, and multi-step agent workflows, combining native text and image understanding with reasoning, tool calling, and structured output. Alibaba recommends it for agentic scenarios such as visual understanding, document and codebase analysis, desktop interaction, and chart interpretation, and the QwenCloud changelog notes full compatibility with both OpenAI and Anthropic API protocols so it can plug into tools like Claude Code and Codex for high-concurrency deployments. Third-party routings such as Vercel's AI Gateway and OpenRouter expose it through unified APIs, with OpenRouter telemetry showing around 53 tokens per second at roughly 3.26 seconds of P50 latency and an 89% average cache hit rate, indicating that the long-context design is paired with infrastructure tuned for production throughput.

SCNet Token PlanQwen3.8-Flashqwen

Quick Info

Powered by
Provider
SCNet Token Plan
Model key
Qwen3.8-Flash
Release date
Aug 26, 2026
Last updated
Aug 26, 2026
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
131,072 tokens
Context window
1,000,000 tokens

Latest news about Qwen3.8 Flash

SCNet Token Plan

CoverageRelease Notes

Vercel announced that Qwen 3.8 Flash from Alibaba is now available on its AI Gateway, taking text and images as input, serving a context window of 1 million tokens, and returning up to 65k tokens per response. Alibaba recommends the model for coding, tool use, and multi-step agent workflows. Developers can invoke it by Vercel's AI Gateway provides a unified API with usage tracking, retry, failover, custom reporting, Zero Data Retention support, and budgets. The gateway reflects provider pricing with no markup and does not charge a platform fee on inference, including on Bring Your Own Key requests. This integration post independently

SCNet Token Plan

Coverage

OpenRouter lists qwen3.8-flash as a multimodal reasoning model from Alibaba, suited for coding assistance, agentic workflows, visual understanding, document and codebase analysis, desktop interaction, chart analysis, and long-video analysis. The listing confirms a release date of August 26, 2026, a 1M context window, a OpenRouter's telemetry reports throughput of 53 tokens per second (P50) and P50 latency of 3.26 seconds for Alibaba Cloud International, along with an average cache hit rate of 89.36%. The weighted-average effective input price is $0.03325 per 1M tokens, reflecting heavy caching, while weighted-average effective output

SCNet Token Plan

CoverageRelease Notes

Alibaba's QwenCloud officially announced the release of qwen3.8-flash on August 26, 2026, as documented in the provider's first-party changelog. The model is described as the latest multimodal addition to the Qwen family, natively supporting a 1M-token context window for processing long documents, entire codebases, and According to the same QwenCloud changelog entry, qwen3.8-flash is fully compatible with both OpenAI and Anthropic API protocols, enabling integration with developer tools such as Claude Code and Codex for high-concurrency applications. This first-party documentation provides the authoritative anchor for the model's gen

Videos about Qwen3.8 Flash

More models around Qwen3.8 Flash