SD 2 Omni Reference VIP 4K by ByteDance. Generate 4K ultra-HD videos using up to 9 image references, up to 3 video clips, and up to 3 audio references with priority routing. Reference materials in your prompt with @image1…@image9, @video1…@video3, and @audio1…@audio3.
About this model
Seedance 2 Omni Reference VIP 4K extends our multi-modal reference system to 4K ultra-HD quality. Combine up to 9 images, 3 video clips, and 3 audio files into a single 4K generation using the same intuitive reference notation (@image1-@image9, @video1-@video3, @audio1-@audio3). The model synthesizes all these materials into a pristine 3840×2160 video with full audio-visual synchronization, generating 4 to 15 second outputs that respect the visual language and audio character of every reference you provide.
Ideal for premium production work where quality, complexity, and creative ambition converge. Professional studios, high-end agencies, and ambitious indie creators use this tier to produce broadcast-grade content that blends multiple sources into seamless, cohesive videos. Muapi's unified API and per-generation pricing let you create at any scale without subscription commitments.
Cost analysis
| Provider | Cost | Notes |
|---|---|---|
| muapiapp | $6.75 per generation | Muapi's transparent per-generation model ensures predictable costs for complex multi-modal 4K productions, with no subscription overhead or hidden reference-complexity surcharges. |
Muapi's transparent per-generation model ensures predictable costs for complex multi-modal 4K productions, with no subscription overhead or hidden reference-complexity surcharges.
* Competitor pricing is estimated based on similar model architectures and usage tiers.
Configuration schema
| Parameter | Type | Description | Default |
|---|---|---|---|
| Prompt | string | Video description. Use @image1…@image9 to reference images, @video1…@video3 for videos, and @audio1…@audio3 for audio. Use @character:<request_id> for a Seedance 2 character sheet or @omni-character:<char_id> for a trained Kinovi character. Multiple characters are supported. | @image1 is the main character. The person walks along a city street at sunset, cinematic lighting. |
| Image URLs | array | Up to 9 reference image URLs (JPEG/PNG/WebP). Each Nth image corresponds to @imageN in the prompt. | https://d3adwkbyhxyrtq.cloudfront.net/ai-images/186/712345784292/4a8c5c70-abcc-4920-873e-b0e219986453.jpg |
| Video Reference URLs | array | Up to 3 reference video clip URLs (MP4, max 15s each). Each Nth video corresponds to @videoN in the prompt. | undefined |
| Audio Reference URLs | array | Up to 3 reference audio files (MP3/WAV, total max 15s). Each Nth audio corresponds to @audioN in the prompt. | undefined |
| Aspect Ratio | Enum (6 options) | Output video aspect ratio. | 16:9 |
| Duration (seconds) | int | Video duration in seconds. | 5 |
Video description. Use @image1…@image9 to reference images, @video1…@video3 for videos, and @audio1…@audio3 for audio. Use @character:<request_id> for a Seedance 2 character sheet or @omni-character:<char_id> for a trained Kinovi character. Multiple characters are supported.
@image1 is the main character. The person walks along a city street at sunset, cinematic lighting.Up to 9 reference image URLs (JPEG/PNG/WebP). Each Nth image corresponds to @imageN in the prompt.
https://d3adwkbyhxyrtq.cloudfront.net/ai-images/186/712345784292/4a8c5c70-abcc-4920-873e-b0e219986453.jpgUp to 3 reference video clip URLs (MP4, max 15s each). Each Nth video corresponds to @videoN in the prompt.
undefinedUp to 3 reference audio files (MP3/WAV, total max 15s). Each Nth audio corresponds to @audioN in the prompt.
undefinedOutput video aspect ratio.
16:9Video duration in seconds.
5Developer documentation
Frequently asked
4K does require more computational resources, but muapi's infrastructure is optimized for efficient 4K generation. Processing time is reasonable for a production tool—typically completing within a couple of minutes depending on complexity and system load.
Standard formats (JPEG/PNG for images, MP4/WebM for video, WAV/MP3 for audio) are supported. For best results, provide high-quality source materials—1080p or higher for images, high-bitrate video, and clear audio tracks. The model will handle the synthesis efficiently.
This model costs $6.75 per generation. The price is fixed regardless of reference complexity—whether you use 3 materials or all 15 maximum, you pay one transparent per-generation rate.