Blend images, video, and audio into one video with Seedance 2.5 Omni Reference. 720p, 30-second clips. Try free — pay per generation, no subscription. Relaxed-moderation (Spicy) endpoint, priced 10% above standard.
About this model
Seedance 2.5 Omni Reference is the early-access multi-reference endpoint in the Seedance 2.5 family, blending up to 30 reference images, 10 video clips, and 10 audio files into a single guided video generation. Use the images to steer environment and style, the videos for camera motion and rhythm, and the audio as a mood reference — all combined via one text prompt. Supports clips up to 30 seconds at 720p. This is Muapi's Spicy endpoint for the same Seedance 2.5 model — the relaxed-moderation sibling of the standard tier, with lighter content-safety filtering and bolder, higher-contrast output.
Cost analysis
| Provider | Cost | Notes |
|---|---|---|
| muapiapp | $0.34/sec (720p) or $0.17/sec (480p) | Rate applies to output duration only when no video reference is used. With a video reference, billed at 65% of that rate ($0.221/sec at 720p, $0.11/sec at 480p) on the combined duration of the reference videos plus the generated output. |
| Fal.ai | $0.47/sec (720p) or $0.2205/sec (480p) | MuAPI is ~28% cheaper than Fal.ai at 720p and ~23% cheaper at 480p. |
| Replicate | Not available | Seedance 2.5 is not listed on Replicate. |
Rate applies to output duration only when no video reference is used. With a video reference, billed at 65% of that rate ($0.221/sec at 720p, $0.11/sec at 480p) on the combined duration of the reference videos plus the generated output.
MuAPI is ~28% cheaper than Fal.ai at 720p and ~23% cheaper at 480p.
Seedance 2.5 is not listed on Replicate.
* Competitor pricing is estimated based on similar model architectures and usage tiers.
Configuration schema
| Parameter | Type | Description | Default |
|---|---|---|---|
| Prompt | string | Text prompt describing the desired video, referencing the provided images, video clips, and audio as environment, motion, and mood cues. | Create a calm cinematic park sequence. Use the images for environment style, the video clips for camera motion and street rhythm, and the audio as mood reference. |
| Reference Images | array | Reference image URLs. Up to 30 images. | undefined |
| Reference Videos | array | Reference video URLs. Up to 10 clips. | undefined |
| Reference Audio | array | Reference audio URLs. Up to 10 files. | undefined |
| Duration | int | The duration of the generated video in seconds. | 5 |
| Aspect Ratio | Enum (7 options) | Aspect ratio of the output video. | 16:9 |
| Seed | int | Random seed for reproducible generation. Use -1 for random. | 42 |
Text prompt describing the desired video, referencing the provided images, video clips, and audio as environment, motion, and mood cues.
Create a calm cinematic park sequence. Use the images for environment style, the video clips for camera motion and street rhythm, and the audio as mood reference.Reference image URLs. Up to 30 images.
undefinedReference video URLs. Up to 10 clips.
undefinedReference audio URLs. Up to 10 files.
undefinedThe duration of the generated video in seconds.
5Aspect ratio of the output video.
16:9Random seed for reproducible generation. Use -1 for random.
42Developer documentation
Gather Your References
images_list, 10 videos via videos_list, 10 audio files via audios_list. At least one reference is recommended.Write a Guiding Prompt
Configure Output
duration (4–30s), aspect_ratio, and optionally seed.Submit via the API
curl -X POST https://api.muapi.ai/api/v1/seedance-2.5-spicy-omni-reference \
-H "x-api-key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"prompt": "Create a calm cinematic park sequence using the references for style, motion, and mood.",
"images_list": ["https://example.com/ref1.jpg", "https://example.com/ref2.jpg"],
"videos_list": ["https://example.com/ref.mp4"],
"duration": 10
}'
GET /api/v1/predictions/{request_id}/result until status is completed, then retrieve the video URL from output.Frequently asked
Same request/response shape as the standard endpoint, but served on the relaxed-moderation Spicy tier, with reduced content-safety filtering and more dramatic, higher-contrast motion. Priced 10% above the standard tier.
It generates video guided by a combination of reference images, video clips, and audio files, blending their style, motion, and mood cues into one output.
Up to 30 images, 10 video clips, and 10 audio files, via `images_list`, `videos_list`, and `audios_list` respectively.
Omni Reference uses the same per-second rate as the base model ($0.34/sec at 720p, $0.17/sec at 480p), but is billed on total video duration — the generated output plus every reference video you supply — not just the output length.
720p by default. A cheaper 480p variant is available separately.
Between 4 and 30 seconds, configurable via the `duration` parameter.
No — provide any combination of images, videos, and audio. A text prompt alone still works, but references help steer style, motion, and mood more precisely.