Currently listed through these providers:
Model details
Wan v2.7 Reference-to-Video
Wan v2.7 Reference-to-Video is a diffusion-based generative video model built for reference-conditioned synthesis, where a text prompt plus a reference still are fused into a moving clip that preserves the subject's identity, wardrobe, and visual style. It sits inside the broader Wan family of video models and is described as the latest generation in that line, with refinements aimed at smoother motion, better scene fidelity, and stronger visual coherence than prior Wan releases. The model is delivered as a video endpoint (a POST to the fal video runtime) that accepts a prompt and an optional reference image and returns rendered video, making it well suited to creative pipelines that need a known visual anchor rather than purely text-driven generation.
In practice the model fits workflows that demand subject consistency across shots, such as product demos, brand vignettes, fashion or character reveals, and short marketing clips where a hero object or person must remain recognizable. Its supported input modalities of text and image, combined with adjustable resolution, duration, and aspect ratio, let creators prototype cinematic compositions from a single reference still while keeping creative control over framing and pacing. Compared with earlier Wan 2.2 variants in the same family, v2.7 Reference-to-Video emphasizes consistency from reference inputs over open-ended text-to-video generation, making it a strong choice when stable identity and on-style visuals matter more than maximal motion range.
Quick Info
Powered by- Provider
- Vercel AI Gateway
- Model key
- alibaba/wan-v2.7-r2v
- Release date
- Apr 7, 2026
- Last updated
- Apr 7, 2026
- Input modalities
- Output modalities
- Capabilities
Limits
- Output tokens
- 0 tokens
- Context window
- 0 tokens