Currently listed through these providers:
Model details
Wan v2.7 Text-to-Video
Wan v2.7 Text-to-Video is Alibaba's latest-generation generative video model, surfaced as a serverless endpoint that translates written prompts directly into finished clips. It is presented as the newest entry in Alibaba's Wan video family, with the marketing positioning centered on smoother motion, sharper scene fidelity, and more coherent temporal behavior than earlier Wan releases. The text-to-video route on fal.ai (fal-ai/wan/v2.7/text-to-video) is the primary endpoint documented in the available evidence, framed under Alibaba branding alongside Qwen visual assets that signal its place within the broader Alibaba model portfolio.
In practical terms, the model is aimed at creators and product teams who need prompt-driven HD video output up to roughly fifteen seconds, including native audio, first-and-last-frame control for shaping scene start and end, instruction-based editing, and character reference for consistent on-screen identity. These features position it as a flexible foundation for short-form content, storyboard prototyping, and visually directed generation where motion quality and continuity matter. Because the available evidence is limited to hosting-side documentation rather than primary Alibaba technical reports, the strongest fit is for projects that value the integrated editing and audio capabilities over independently benchmarked performance claims.
Quick Info
Powered by- Provider
- Vercel AI Gateway
- Model key
- alibaba/wan-v2.7-t2v
- Release date
- Apr 7, 2026
- Last updated
- Apr 7, 2026
- Input modalities
- Output modalities
- Capabilities
Limits
- Output tokens
- 0 tokens
- Context window
- 0 tokens