Vidu Q2 Reference Video generates breathtaking cinematic clips from text prompts guided by multiple reference images. Each image refines the model’s understanding of subject, environment, and visual tone — ensuring perfect consistency in appearance and motion across every frame.
About this model
Vidu Q2 Reference Video is a state-of-the-art image-to-video generation model that transforms text prompts and multiple reference images into breathtaking cinematic clips. Leveraging advanced deep learning techniques and sophisticated image processing technology, it meticulously refines each frame’s subject, environment, and visual tone to ensure perfect consistency in appearance and motion. The model’s ability to merge detailed textual descriptions with visual references sets a new benchmark for creative video production.
This robust technology is not only capable of generating high-quality videos at resolutions up to 1080p, but it also offers customizable parameters such as aspect ratio, duration, and movement amplitude. Whether used for professional filmmaking, advertising, or social media content creation, Vidu Q2 Reference Video gives creators unparalleled control and flexibility, enabling them to bring their artistic visions to life with cinematic precision and flair.
Cost analysis
| Provider | Cost | Notes |
|---|---|---|
| muapiapp | $0.065 per generation | Offers exceptional quality and is 20-50% more affordable than leading competitors. |
| Fal.ai | $0.10 per generation | Priced higher than muapiapp, making muapiapp a more cost-effective choice with comparable or superior quality. |
| Replicate | $0.10 per generation | Similarly priced to Fal.ai. Muapiapp delivers a 20-50% cost saving while matching their performance and quality. |
Offers exceptional quality and is 20-50% more affordable than leading competitors.
Priced higher than muapiapp, making muapiapp a more cost-effective choice with comparable or superior quality.
Similarly priced to Fal.ai. Muapiapp delivers a 20-50% cost saving while matching their performance and quality.
* Competitor pricing is estimated based on similar model architectures and usage tiers.
Configuration schema
| Parameter | Type | Description | Default |
|---|---|---|---|
| Prompt | string | The prompt to generate the video | The female explorer walks slowly across the alien terrain, crystals glimmering around her. The camera glides beside her as light from twin suns scatters across her reflective suit. Wind stirs the mist as she looks up toward the horizon, where a colossal planet looms above — evoking awe and wonder. |
| Image URLs | array | Upload or provide image urls. Used for image-to-video generation. | https://d3adwkbyhxyrtq.cloudfront.net/webassets/videomodels/vidu-q2-reference-1.jpg |
| Resolution | Enum (4 options) | The resolution of the generated video. | 720p |
| Aspect Ratio | Enum (5 options) | Aspect ratio of the output video. | 16:9 |
| Duration | int | The duration of the generated video in seconds | 5 |
| Movement Amplitude | Enum (4 options) | The movement amplitude of objects in the frame. | auto |
The prompt to generate the video
The female explorer walks slowly across the alien terrain, crystals glimmering around her. The camera glides beside her as light from twin suns scatters across her reflective suit. Wind stirs the mist as she looks up toward the horizon, where a colossal planet looms above — evoking awe and wonder.Upload or provide image urls. Used for image-to-video generation.
https://d3adwkbyhxyrtq.cloudfront.net/webassets/videomodels/vidu-q2-reference-1.jpgThe resolution of the generated video.
720pAspect ratio of the output video.
16:9The duration of the generated video in seconds
5The movement amplitude of objects in the frame.
autoDeveloper documentation
Prepare Your Inputs
Set Your Parameters
Generate and Review
Download and Share
Frequently asked
The model supports multiple resolutions, including 360p, 540p, 720p (default), and 1080p, allowing you to choose the best quality for your needs.
The provided reference images help guide the model by refining details like subject appearance, environmental elements, and overall visual tone. This ensures consistency in motion and appearance throughout every frame of the video.
You can upload up to 7 reference images to guide the video generation process.
The model allows you to adjust several parameters including resolution, aspect ratio, duration, and movement amplitude, giving you full control over the cinematic output.