Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
NaN logo

Model details

Qwen-Image-2.1

Alibaba's Qwen team released Qwen-Image-2.1 as an open-weights image generation and editing model that continues the line established by its predecessor, sharing an optimized MMDiT-style backbone. At 7B parameters, the model is notably compact for its class, allowing the weights to run comfortably on consumer hardware while still producing native 2K-resolution imagery. The open-weight release invites local fine-tuning and integration into custom pipelines rather than locking users into a hosted service.

Qwen-Image-2.1 is designed around practical production and creative workflows, generating a real alpha channel for true RGBA output that eliminates the need for separate background-removal or matting passes when compositing sprites, logos, icons, or product cutouts. It also accepts up to ten input images simultaneously, enabling coherent multi-reference editing sessions for tasks like style transfer, character consistency, and layout rearrangement. Native ComfyUI support broadens accessibility for hobbyists and studios seeking a transparent, locally deployable image model with strong editing versatility.

NaNqwen-image-2.1qwen

Quick Info

Powered by
Provider
NaN
Model key
qwen-image-2.1
Release date
Sep 14, 2026
Last updated
Sep 21, 2026
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
0 tokens
Context window
8,192 tokens

Latest news about Qwen-Image-2.1

NaN

CoverageBenchmark

Tom's Hardware reports that Alibaba's Qwen Image 2.1, a 7B-parameter open-weight image model, is being benchmarked as competitive with OpenAI and Meta image offerings and claimed by Alibaba to outperform Google's Nano Banana 2.0. The article frames the release as a compact open-weight challenger in the generative image space, leveraging a much smaller parameter count than the prior 20B Qwen image model. The coverage positions Qwen-Image 2.1 within the broader competitive landscape of image generation models, with benchmark framing against proprietary systems from Google, OpenAI, and Meta. This adds competitive-context perspective to the model's open-weight release, though the superiority claims trace back to Alibaba-provided benchmarks rather than independent verification.

NaN

Coverage

Alibaba's Qwen team released Qwen-Image 2.1 on September 20, 2026 via Hugging Face, ModelScope, and GitHub, combining text-to-image generation and image editing in a single model. The image generator is a 7B-parameter diffusion transformer built from 32 Single-Stream DiT layers, a sharp shrink from the prior 20B model that needed 40.9GB BF16; it lands at 14.2GB BF16 or 7.26GB INT8, making it runnable on consumer hardware while keeping most prior capabilities. The headline new feature is native RGBA four-channel output: Qwen-Image 2.1 decides from the prompt whether to emit a normal image or a transparent PNG and generates the alpha channel directly, removing the usual generate-then-cutout post-processing step. Day-zero ecosystem support shipped in ComfyUI, Diffusers, vLLM-Omni, and SGLang, alongside a native 2K output size and support for up to ten input images for multi-region editing.

NaN

Coverage

Intel announced day-zero OpenVINO support for Qwen-Image 2.1 ahead of the Apsara Conference, optimizing the 7B-parameter model for Intel Arc Pro B70 GPUs and Intel Core Ultra Series 3 integrated GPUs. This gives developers a local deployment path with GPU-accelerated inference on Intel hardware from the release date, confirming the model's "ultra-fast inference" positioning. Intel's coverage corroborates Qwen-Image 2.1's native 2K image output, support for up to ten input images, and native RGBA four-channel transparent generation and editing without post-processing. Combined with multi-region and multi-type local editing, the integration positions the model for both creative and professional workflows on Intel-accelerated systems.

NaN

Coverage

Qwen has open-sourced Qwen-Image-2.1, a compact image model unifying text-to-image generation and image editing in a single system. The visual generation component uses 7B parameters across 32 Single-Stream DiT layers, delivering quality comparable to larger open-source and closed-source models per Qwen-Image-Bench results. Key improvements include native transparent image generation and editing, a mixed-granularity attention architecture with token-level causal masking for text and chunk-level masking for images plus KV cache reuse for efficiency, and support for up to 10 reference images with local edit specification, improved typography, portrait lighting, and fine detail rendering.

NaN

Coverage

Hacker News community testing confirms Qwen-Image-2.1 is significantly smaller than Qwen-Image 1.0 (7B vs 20B parameters) while being faster and more coherent at 2K output resolutions. Third-party GenAI Showdown benchmarks scored the model 7 out of 15, a notable leap from its predecessor's 4 out of 15 score. Community testers reported inference speeds of approximately 5 seconds per 1MP image on an RTX 4090 using QwenImage2.1 convrot, and noted native transparency support as a distinguishing feature. One user flagged that the license is more restrictive than the original Apache license, explicitly requiring a separate license for commercial use.

Videos about Qwen-Image-2.1

More models around Qwen-Image-2.1