Currently listed through these providers:
Model details
Grok Imagine Video
Grok Imagine Video is the everyday workhorse in xAI's broader "Imagine" multimodal suite, designed to turn images, text, and reference clips into finished video without the patchwork of separate animation, editing, and post-processing tools. It is positioned as a versatile, production-ready generator that prioritizes rapid iteration and predictable cost, making it well-suited to high-volume creative pipelines where many variations need to be tested quickly. The model shares its interface with sibling video models, allowing developers to mix and match jobs like image-to-video, reference-based creation, extension, and editing through a single API surface rather than stitching together several third-party services.
The Imagine Video line is part of xAI's effort to unify image, editing, and video generation under one platform, and it has been refined iteratively toward richer, more cinematic output. More recent iterations in the family have pushed toward coherent long-clip motion, believable weight and momentum, and audio generated in the same pass as the visuals, so sound effects, ambience, and dialogue land on the action with clear, well-synced speech. A dedicated fast tier pushes 6-second 720p clips out in roughly 25 seconds, halving the wait of the prior generation, which makes Grok Imagine Video a practical choice for creators who need quick turnarounds alongside a more computationally intensive option for higher-fidelity work within the same ecosystem.
Quick Info
Powered by- Provider
- xAI
- Model key
- grok-imagine-video
- Release date
- Jan 28, 2026
- Last updated
- Jan 28, 2026
- Input modalities
- Output modalities
- Capabilities
Limits
- Output tokens
- 0 tokens
- Context window
- 1,024 tokens