Currently listed through these providers:
Model details
Qwen 3.7 Flash
Positioned as the efficiency tier of the Qwen 3.7 family from Alibaba, Qwen 3.7 Flash is designed for latency-sensitive and budget-conscious agent pipelines, where predictable cost and quick response matter more than maximum depth. It advertises a roughly 1M token context window (cataloged at the cataloged API limit by pi.dev) with a 64,000-token output ceiling, an optional thinking mode for reasoning-heavy calls, and native tool calling, giving agent frameworks room to attach functions, file references, and structured I/O without bespoke plumbing. Inputs cover text and image, and the gateway exposes the model behind an Anthropic-messages-compatible endpoint at https://ai-gateway.vercel.sh, so existing Anthropic-style client code can route traffic without modification.
Pricing is tiered by total input tokens per request, with the cheapest band applying to short-context workloads and rates scaling up past the 32K and 256K thresholds; pi.dev reports a flat headline of $the listed price input / $the listed price output per million tokens plus $the listed price cache read and $the listed price cache write, while NetMind's ≤32K tier quotes $0.0274 input, $0.1096 output, and $0.00548 cache read. In practice this shape makes Qwen 3.7 Flash a natural fit for high-volume agent loops, long-tail RAG over very large documents, and cost-disciplined batch jobs where the 1M context can be amortized across many cached reads, rather than for frontier reasoning benchmarks or tasks that demand top-of-class accuracy.
Quick Info
Powered by- Provider
- Vercel AI Gateway
- Model key
- alibaba/qwen3.7-flash
- Release date
- Jul 15, 2026
- Last updated
- Jul 15, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.03
- Output token cost
- $0.13
Limits
- Output tokens
- 64,000 tokens
- Context window
- 991,000 tokens
Transparent token rates
Compare Qwen 3.7 Flash pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.