Sulat.com
AI models
Alibaba Token Plan (China) logo

Model details

Qwen3.7 Plus

Qwen3.7 Plus sits in Alibaba's Qwen3.7 family as the vision-enabled sibling of the flagship text model. Built on top of the Qwen3.7 text backbone, it adds native image and video perception so that vision and language are processed jointly from the earliest layers of the network, rather than as a bolt-on captioning module. The result is a perception-only model: it ingests text, images, and short video clips, and replies in text, making it well suited to understanding screens, reasoning about scenes, reading diagrams, and turning visual references such as mockups or screenshots into executable code. Its design intent is what Alibaba calls a multimodal interactive hybrid agent, meaning it can ground itself in graphical interfaces, navigate mobile apps end to end, and combine that perception with coding, reasoning, and tool-calling skills inherited from the broader Qwen3.7 line.

Because Qwen3.7 Plus extends an existing text backbone rather than standing alone, it carries forward the family strengths in coding, long-running reasoning, and structured tool use, while expanding the input surface to include video and image tokens alongside text. Early-fusion training means vision and language representations are learned together, which is what enables GUI grounding and direct code generation from visual references instead of relying on a separate OCR step. Practically, the model is positioned for agentic productivity work: reading a UI, clicking through it, answering questions about a short clip, and writing the code that ties those actions together, all within an extended context window that supports long documents and multi-step tasks. It is best understood as the perception half of a larger agent stack, where it handles the seeing and reasoning while other Alibaba models handle pure text or generative media workloads.

Alibaba Token Plan (China)qwen3.7-plusqwen

Quick Info

Powered by
Provider
Alibaba Token Plan (China)
Model key
qwen3.7-plus
Release date
Jun 2, 2026
Last updated
Jun 2, 2026
Knowledge cutoff
2025-04
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
65,536 tokens
Context window
1,000,000 tokens

Latest news about Qwen3.7 Plus

Alibaba Token Plan (China)

Coverage

MLQ News reports that Alibaba's Qwen team released Qwen3.7-Plus on June 2 as a multimodal agent model that accepts text, images, and video as input but outputs only text, combining visual perception, GUI control, and code generation in a single autonomous loop. The article claims the model leads Alibaba's results on Sc The report notes Alibaba's Hong Kong-listed shares rose as much as 6.84% on the announcement before giving back gains. It frames Qwen3.7-Plus as a "perception and action" model designed to read screens, navigate apps, write code from visual templates, and invoke external tools without human intervention, though it trai

Alibaba Token Plan (China)

CoverageBenchmark

BenchLM.ai's catalog entry for Qwen3.7 Plus, released June 3, 2026, reports a proprietary reasoning model with 1M-token context, an overall score of 62.3 out of 100, and a rank of 57 of 232 models. Capability percentiles show Coding at the 73rd, Knowledge at the 73rd, Multilingual at the 82nd, Multimodal at the 64th, a The catalog notes no comparable first-party API token rate is published and that some tracked benchmark slots remain empty, framing the 52 published benchmark rows as a partial view. Capability ranks include Agentic at 126 of 151 and Multimodal at 18 of 48, providing independent aggregator-level coverage of Qwen3.7-Plu

Alibaba Token Plan (China)

CoverageBenchmark

The llm-stats.com benchmark catalog provides model-specific data for Qwen3.7-Plus, ranking it 43rd overall on its composite LLM Stats Score. Reported benchmark scores include IFEval at 0.95 (rank 2), MMLU-Redux at 0.94 (rank 3), HMMT Feb 26 at 0.93 (rank 4), and MRCR v2 at 0.92 (rank 1), with sources attributed to qwen The page lists category rankings including Chat (6 of 125), Long Context (10 of 114), Math (17 of 317), Vision (20 of 203), Tool Calling (37 of 183), Reasoning (38 of 349), and Coding (46 of 257). However, the page does not display a comparable first-party API price, and benchmark provenance relies on the model's own s

Videos about Qwen3.7 Plus

More models around Qwen3.7 Plus