Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
ZenMux logo

Model details

Qwen3.5 Flash

Qwen3.5 Flash is part of Alibaba's Qwen3.5 release and is positioned as a vision-language Flash-tier model built on a hybrid architecture that combines a linear attention mechanism with a sparse mixture-of-experts design. According to the model listing, this architectural blend is intended to deliver higher inference efficiency, and the Qwen3.5 generation is described as a meaningful performance leap over the prior Qwen3 series for both pure text and multimodal workloads, while keeping response times fast and balancing speed against overall quality. The Flash variant itself is reported to be proprietary rather than openly released, sitting alongside the open-weight Qwen3.5-35B-A3B, Qwen3.5-122B-A10B, and Qwen3.5-27B models that Alibaba shipped under an Apache 2.0 license.

In practical terms, Qwen3.5 Flash is aimed at developers who want the Qwen3.5 family's quality and long-context handling without deploying the larger open-weight variants. The catalog entry indicates support for reasoning and agent-style tool calling, structured output, and adjustable temperature control, with a very wide context window suitable for document-heavy or multi-step agentic flows. Because it is a Flash-tier release, the design priorities emphasize fast inference and cost-effective serving, making it a reasonable fit for high-throughput applications, retrieval-augmented pipelines, and agent workflows where low latency and the Qwen3.5 architecture's efficiency gains matter more than running the largest open checkpoints.

ZenMuxqwen/qwen3.5-flash

Quick Info

Powered by
Provider
ZenMux
Model key
qwen/qwen3.5-flash
Release date
Mar 20, 2026
Last updated
Mar 20, 2026
Knowledge cutoff
2025-01-01
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.10
Output token cost
$0.40

Limits

Output tokens
1,020,000 tokens
Context window
1,020,000 tokens

Latest news about Qwen3.5 Flash

Merge Gateway

Coverage

The QwenCloud model marketplace page documents Qwen3.5-Flash, describing it as a native vision-language Flash model built on a hybrid architecture that combines a linear attention mechanism with a sparse mixture-of-experts model to deliver higher inference efficiency. It supports text, image, and video inputs and text The page lists supported features including prefix completion, function calling, context caching, structured outputs, batch processing, web search, and fine-tuning. Pricing is listed at $0.10 per 1M input tokens and $0.40 per 1M output tokens, with rate limits of up to 1M context, 991K max input, 65K max output, 5M TPM

Videos about Qwen3.5 Flash