Currently listed through these providers:
Model details
Wan v3.0 Video
Wan v3.0 Video is a single, integrated video generation system that handles several creative entry points in one pipeline. It can start from a text prompt, animate a still image, continue from first and last key frames, or assemble output from a mix of image, video, and audio references. The same model is responsible for producing synchronized audio alongside the visuals, so a generated clip arrives with both picture and sound rather than as silent footage. Because it accepts multiple reference types simultaneously, it is suited to tasks where the creator wants fine control over character, scene, or motion cues that cannot be expressed in text alone.
Practically, the model is exposed through a developer-friendly generateVideo interface that accepts a prompt and returns a finished video, with the platform handling routing and waiting for the clip to finish rendering. Each generation can run up to about 30 seconds of video per call, which is long enough for short-form social content, previews, and concept pieces. Input attachments include up to ten images under roughly 20 MB each and optional reference videos under 100 MB, giving creators room to combine several visual references in one request. The combination of omni-modal references, built-in audio, and a polled asynchronous API makes it a flexible choice for builders who want one model to cover ideation, reference-driven generation, and final short-clip production without stitching together separate tools.
Quick Info
Powered by- Provider
- Vercel AI Gateway
- Model key
- alibaba/wan-v3.0-video
- Release date
- Aug 23, 2026
- Last updated
- Aug 23, 2026
- Input modalities
- Output modalities
- Capabilities
Limits
- Output tokens
- 0 tokens
- Context window
- 0 tokens