Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Vercel AI Gateway logo

Model details

Qwen3 Max

Qwen3 Max is Alibaba Cloud's flagship large language model from the Qwen3 family, positioned as a trillion-parameter system intended for advanced reasoning and long-context workloads. An early-access preview of the model is offered through gateway and reseller listings, giving developers ahead-of-schedule access to the underlying Qwen3-Max capabilities for evaluation and prototyping before wider availability.

The model operates as a text-in, text-out system with a 262,144-token shared context window and a maximum output of 32,768 tokens per request. It supports tool calling and temperature control, making it suitable for agent-style applications that need structured function invocation alongside configurable sampling. Gateway compatibility flags indicate eager tool input streaming, long cache retention, and cache control on tools, which point to design priorities around sustained multi-turn agent sessions and efficient reuse of large prompt contexts.

Vercel AI Gatewayalibaba/qwen3-maxqwen

Quick Info

Powered by
Provider
Vercel AI Gateway
Model key
alibaba/qwen3-max
Release date
Sep 23, 2025
Last updated
Sep 23, 2025
Knowledge cutoff
2025-04
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$1.20
Output token cost
$6.00

Limits

Output tokens
32,768 tokens
Context window
262,144 tokens

Transparent token rates

Compare Qwen3 Max pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3 Max

Vercel AI Gateway

CoveragePreview

Pi's model card for `alibaba/qwen3-max-preview` via Vercel AI Gateway corroborates the first-party data with a 262,144-token context window, 32,768 max output tokens, and a tiered pricing structure of $1.20 per 1M input tokens and $6 per 1M output tokens, plus a $0.24 cache-read rate. The API surface is listed as Anthr Effective compatibility flags include EagerToolInputStreaming, LongCacheRetention, CacheControlOnTools, and Temperature support. Pi's session-cost estimator illustrates a typical agent task at roughly 25 requests with ~5k starting context totaling ~$0.235 effective cost versus ~$0.728 without caching, based on observed

Vercel AI Gateway

Official sourcePreview

Vercel AI Gateway's status page for Qwen3 Max Preview describes the model as Alibaba Cloud's early-access, trillion-parameter Qwen3-Max release, giving developers ahead-of-schedule access for evaluation and prototyping. Pricing is set at $1.20 per 1M input tokens and $6 per 1M output tokens. The page surfaces a 24-hour Use of the model through the gateway is subject to Alibaba Cloud's Terms and Privacy Policies. The status page mirrors the API documentation by presenting the same model description, pricing, and SDK/API access links. It also provides a direct path to obtaining a gateway API key and reading the developer documentation.

Vercel AI Gateway

Official sourcePreview

Vercel AI Gateway lists Alibaba's Qwen3 Max as the Preview variant under the model ID `alibaba/qwen3-max-preview`, per its first-party API documentation page. The model shares a 262K-token context window between prompt and response and supports a hard cap of 32,768 output tokens. Developers can route requests through t The page documents top-level parameters including the required `model` string in `creator/model` form and an optional `maxOutputTokens` numeric cap, plus `providerOptions` for gateway routing and provider-native settings. Quickstart instructions cover installing the AI SDK, creating an API key, and setting `AI_GATEWAY_

Videos about Qwen3 Max

More models around Qwen3 Max