SD 2 Omni Reference VIP 1080p by ByteDance. Generate full HD videos using up to 9 image references, up to 3 video clips, and up to 3 audio references with priority routing. Reference materials in your prompt with @image1…@image9, @video1…@video3, and @audio1…@audio3.
About this model
Seedance 2 Omni Reference VIP 1080p is a multi-modal video generation model that lets you combine up to 9 images, 3 video clips, and 3 audio files as references in a single generation. Instead of animating a single image, you construct a rich prompt that references these materials using @image1 through @image9, @video1 through @video3, and @audio1 through @audio3 notation. The model synthesizes all these inputs to produce a cohesive 1080p video that respects the visual and audio characteristics you've provided, with full native audio-visual synchronization across 4 to 15 seconds.
This is the go-to choice for complex creative projects: editing existing footage with visual references, layering multiple audio tracks, or blending style references from multiple sources. Access it through muapi's unified API with no subscriptions—only pay when you generate.
Cost analysis
| Provider | Cost | Notes |
|---|---|---|
| muapiapp | $3.38 per generation | Muapi charges a flat per-generation rate regardless of how many reference materials you include, making multi-modal video creation affordable and predictable. |
Muapi charges a flat per-generation rate regardless of how many reference materials you include, making multi-modal video creation affordable and predictable.
* Competitor pricing is estimated based on similar model architectures and usage tiers.
Configuration schema
| Parameter | Type | Description | Default |
|---|---|---|---|
| Prompt | string | Video description. Use @image1…@image9 to reference images, @video1…@video3 for videos, and @audio1…@audio3 for audio. Use @character:<request_id> for a Seedance 2 character sheet or @omni-character:<char_id> for a trained Kinovi character. Multiple characters are supported. | @image1 is the main character. The person walks along a city street at sunset, cinematic lighting. |
| Image URLs | array | Up to 9 reference image URLs (JPEG/PNG/WebP). Each Nth image corresponds to @imageN in the prompt. | https://d3adwkbyhxyrtq.cloudfront.net/ai-images/186/712345784292/4a8c5c70-abcc-4920-873e-b0e219986453.jpg |
| Video Reference URLs | array | Up to 3 reference video clip URLs (MP4, max 15s each). Each Nth video corresponds to @videoN in the prompt. | undefined |
| Audio Reference URLs | array | Up to 3 reference audio files (MP3/WAV, total max 15s). Each Nth audio corresponds to @audioN in the prompt. | undefined |
| Aspect Ratio | Enum (6 options) | Output video aspect ratio. | 16:9 |
| Duration (seconds) | int | Video duration in seconds. | 5 |
Video description. Use @image1…@image9 to reference images, @video1…@video3 for videos, and @audio1…@audio3 for audio. Use @character:<request_id> for a Seedance 2 character sheet or @omni-character:<char_id> for a trained Kinovi character. Multiple characters are supported.
@image1 is the main character. The person walks along a city street at sunset, cinematic lighting.Up to 9 reference image URLs (JPEG/PNG/WebP). Each Nth image corresponds to @imageN in the prompt.
https://d3adwkbyhxyrtq.cloudfront.net/ai-images/186/712345784292/4a8c5c70-abcc-4920-873e-b0e219986453.jpgUp to 3 reference video clip URLs (MP4, max 15s each). Each Nth video corresponds to @videoN in the prompt.
undefinedUp to 3 reference audio files (MP3/WAV, total max 15s). Each Nth audio corresponds to @audioN in the prompt.
undefinedOutput video aspect ratio.
16:9Video duration in seconds.
5Developer documentation
Frequently asked
Yes, you can provide all materials simultaneously, and the model will incorporate them according to your prompt instructions. However, clarity in your prompt matters—specific directions about which reference should be dominant or how they should interact will produce better results than providing references without clear guidance.
Use @image1 through @image9 for images, @video1 through @video3 for videos, and @audio1 through @audio3 for audio clips. For example: '@image1 is the hero shot, @video1 provides background movement, @audio1 is the voiceover.' Reference only the files you've actually uploaded.
This model costs $3.38 per generation. Despite the added complexity of handling multiple reference types, pricing remains straightforward and transparent—no subscriptions, no per-reference surcharges, just a single per-generation fee.