Currently listed through these providers:
Model details
Qwen3.7 Plus
Qwen3.7 Plus is built on the Qwen 3.7 text backbone and represents a focused upgrade toward unified vision-language reasoning. Official Qwen materials describe it as a multimodal agent foundation that combines visual perception with text understanding in a single model, retaining the family’s established strengths in coding, tool use, and productivity workflows. This lineage matters in practice: rather than bolting a vision encoder onto a text model, the design treats visual input as a first-class signal for planning and action, which is why the same checkpoint can answer questions about an image, read on-screen interfaces, and continue with downstream reasoning in one agent loop.
The practical identity of Qwen3.7 Plus is as a multimodal interactive hybrid agent. It can perceive real-world scenes through image and video inputs, read screens and operate GUIs, generate code from visual references, navigate mobile applications end-to-end, and ground visual questions in retrieved web knowledge, all while remaining usable across popular agent scaffolds. Independent third-party listings echo this positioning, describing a mid-to-high tier Plus variant that keeps full agent capabilities rather than trading them away for stronger perception. The result is a model well suited to teams that want one checkpoint to handle mixed-modality inputs and to drive real action, from front-end prototyping and software engineering to multi-step workflow automation that touches both visual context and structured tool calls.
Quick Info
Powered by- Provider
- Pioneer
- Model key
- qwen3.7-plus
- Release date
- Jun 2, 2026
- Last updated
- Jun 2, 2026
- Knowledge cutoff
- 2025-04
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.32
- Output token cost
- $1.28
Limits
- Output tokens
- 64,000 tokens
- Context window
- 1,000,000 tokens