SD 2 Omni Reference VIP 1080p Fast by ByteDance. Faster full HD omni-reference video generation with priority routing. Supports image, video, and audio references.
About this model
Seedance 2 Omni Reference VIP 1080p Fast combines the speed optimization of our Fast tier with the multi-modal reference capability of Omni Reference. Generate 1080p videos using up to 9 images, 3 video clips, and 3 audio files as references—with accelerated processing time. This is the ideal choice for workflows where you need complex multi-modal synthesis at production speed: rapid iteration in creative tools, interactive demos, batch processing multiple variations, or user-facing applications that generate videos on-demand with multiple input types.
You reference materials using the same @image/@video/@audio notation, describe the synthesis in your prompt, and muapi handles the rest. Pay only per generation—no subscriptions, no monthly minimums, just fast, cost-effective video generation at scale.
Cost analysis
| Provider | Cost | Notes |
|---|---|---|
| muapiapp | $2.36 per generation | Muapi's fast-tier pricing makes complex multi-modal synthesis accessible at a lower cost than standard reference tiers, perfect for high-volume creative workflows and batch processing. |
Muapi's fast-tier pricing makes complex multi-modal synthesis accessible at a lower cost than standard reference tiers, perfect for high-volume creative workflows and batch processing.
* Competitor pricing is estimated based on similar model architectures and usage tiers.
Configuration schema
| Parameter | Type | Description | Default |
|---|---|---|---|
| Prompt | string | Video description. Use @image1…@image9 to reference images, @video1…@video3 for videos, and @audio1…@audio3 for audio. Use @character:<request_id> for a Seedance 2 character sheet or @omni-character:<char_id> for a trained Kinovi character. Multiple characters are supported. | @image1 is the main character. The person walks along a city street at sunset, cinematic lighting. |
| Image URLs | array | Up to 9 reference image URLs (JPEG/PNG/WebP). Each Nth image corresponds to @imageN in the prompt. | https://d3adwkbyhxyrtq.cloudfront.net/ai-images/186/712345784292/4a8c5c70-abcc-4920-873e-b0e219986453.jpg |
| Video Reference URLs | array | Up to 3 reference video clip URLs (MP4, max 15s each). Each Nth video corresponds to @videoN in the prompt. | undefined |
| Audio Reference URLs | array | Up to 3 reference audio files (MP3/WAV, total max 15s). Each Nth audio corresponds to @audioN in the prompt. | undefined |
| Aspect Ratio | Enum (6 options) | Output video aspect ratio. | 16:9 |
| Duration (seconds) | int | Video duration in seconds. | 5 |
| High Bitrate | boolean | Enable high bitrate mode for better visual fidelity. Produces larger files. | false |
Video description. Use @image1…@image9 to reference images, @video1…@video3 for videos, and @audio1…@audio3 for audio. Use @character:<request_id> for a Seedance 2 character sheet or @omni-character:<char_id> for a trained Kinovi character. Multiple characters are supported.
@image1 is the main character. The person walks along a city street at sunset, cinematic lighting.Up to 9 reference image URLs (JPEG/PNG/WebP). Each Nth image corresponds to @imageN in the prompt.
https://d3adwkbyhxyrtq.cloudfront.net/ai-images/186/712345784292/4a8c5c70-abcc-4920-873e-b0e219986453.jpgUp to 3 reference video clip URLs (MP4, max 15s each). Each Nth video corresponds to @videoN in the prompt.
undefinedUp to 3 reference audio files (MP3/WAV, total max 15s). Each Nth audio corresponds to @audioN in the prompt.
undefinedOutput video aspect ratio.
16:9Video duration in seconds.
5Enable high bitrate mode for better visual fidelity. Produces larger files.
falseDeveloper documentation
Frequently asked
No—the Fast variant maintains full synthesis quality while accelerating processing. It uses the same advanced multi-modal algorithms as the standard Omni Reference tier; only the generation pipeline is optimized for speed.
You can use up to 15 total references: 9 images, 3 videos, and 3 audio clips. You don't need to max out—provide exactly what your creative vision requires, and let the model work with whatever subset you upload.
This model costs $2.36 per generation. It's the most affordable tier for multi-modal video synthesis, combining speed, flexibility, and cost-efficiency in one offering.