VIDU Reference-to-Image Q2 generates new high-quality images based on one or more reference images. It preserves the key identity, structure, or style of the reference while creating a new scene, variation, or enhanced composition. Ideal for character consistency, object re-interpretation, stylized redesigns, and cinematic recreations guided by reference inputs.
About this model
VIDU Reference-to-Image Q2 is an advanced image-to-image generation model designed to transform one or more reference images into entirely new compositions. Leveraging state-of-the-art AI and deep learning techniques, this model maintains the key identity, structure, and style of input images while creating enhanced variations, fresh scenes, or cinematic recreations. The underlying technology ensures that intricate details and artistic nuances from the reference images are preserved, making it highly effective for applications that require consistency and creativity.
In addition to its robust technical foundation, VIDU Reference-to-Image Q2 offers significant marketing advantages. Its ability to generate high-quality and stylistically coherent images from multiple inputs supports diverse use cases, from character consistency in entertainment to object reinterpretation in product design. Its competitive cost of $0.032 per generation further reinforces its appeal, providing superior value without compromising on output quality.
Cost analysis
| Provider | Cost | Notes |
|---|---|---|
| muapiapp | $0.032 per generation | muapiapp is 20-50% more affordable than competitors while delivering comparable or superior output quality. |
| Fal.ai | $0.045 per generation | Fal.ai charges close to this price point, making muapiapp a cost-efficient alternative with similar high-quality results. |
| Replicate | $0.045 per generation | Replicate's pricing is nearly identical to Fal.ai, meaning muapiapp offers a 20-50% cost saving compared to these providers. |
muapiapp is 20-50% more affordable than competitors while delivering comparable or superior output quality.
Fal.ai charges close to this price point, making muapiapp a cost-efficient alternative with similar high-quality results.
Replicate's pricing is nearly identical to Fal.ai, meaning muapiapp offers a 20-50% cost saving compared to these providers.
* Competitor pricing is estimated based on similar model architectures and usage tiers.
Configuration schema
| Parameter | Type | Description | Default |
|---|---|---|---|
| Prompt | string | Text prompt describing the image. | Create a new scene where the masked wanderer stands inside an ancient stone observatory illuminated by rotating celestial beams; preserve the character’s clothing style and silhouette while adding glowing runes carved into the walls, mist swirling across the floor, and a dramatic cosmic light shaft from above; cinematic composition, high detail. |
| Image URLs | array | Upload or provide reference images. Used for image-to-image generation. | https://d3adwkbyhxyrtq.cloudfront.net/webassets/videomodels/vidu-q2-reference-to-image-in.jpg |
| Aspect Ratio | Enum (9 options) | Aspect ratio of the output image. | 1:1 |
| Resolution | Enum (3 options) | The target resolution of the generated image. | 1k |
Text prompt describing the image.
Create a new scene where the masked wanderer stands inside an ancient stone observatory illuminated by rotating celestial beams; preserve the character’s clothing style and silhouette while adding glowing runes carved into the walls, mist swirling across the floor, and a dramatic cosmic light shaft from above; cinematic composition, high detail.Upload or provide reference images. Used for image-to-image generation.
https://d3adwkbyhxyrtq.cloudfront.net/webassets/videomodels/vidu-q2-reference-to-image-in.jpgAspect ratio of the output image.
1:1The target resolution of the generated image.
1kDeveloper documentation
Prepare Your Inputs
Configure the Technical Settings
Initiate Generation
vidu-q2-reference-to-image).Review and Interpret Results
Frequently asked
The model uses advanced deep learning techniques to capture and preserve key characteristics such as structure and style from the reference inputs, ensuring that the generated image reflects the core identity of the original.
Detailed and descriptive text prompts yield the best results. Including specific directions regarding style, composition, and desired variations helps the model align the final output with your vision.
Yes, you can provide up to 7 reference images. This allows you to combine multiple sources of inspiration while ensuring the model can process them effectively.
You can choose from 1k, 2k, or 4k resolutions. The selection depends on your project's quality requirements and output medium.