Seedance 2 全能参考: Image-to-Video

SD 2.0 全能参考可使用参考图像、视频和音频生成具有视觉一致性的视频。保持角色身份、视觉风格和场景连续性。单次请求最多支持 9 张图像、3 个视频片段和 3 个音频文件。可在Prompt中使用 @image1、@video1、@audio1 语法,精确控制每个参考素材对生成视频的影响。

📝

Overview

About this model

SD 2.0 全能参考可使用参考图像、视频和音频生成具有视觉一致性的视频。不同于只将单张图像制作成动画的标准Imagem para Vídeo,全能参考会将你上传的素材用作创作引导——在保持角色身份、视觉风格和场景连续性的同时生成新视频。单次请求可组合最多 9 张图像、3 个视频片段和 3 个音频文件。在Prompt中使用 @image1、@video1、@audio1 语法,可精确控制每个参考素材如何影响生成的视频。

1角色一致性:将肖像作为 @image1,保持角色在多个场景中的外观一致。
2风格迁移:将参考图像的视觉风格应用到新生成的视频场景。
3音频同步视频:通过 @audio1 生成与参考音乐片段或语音录音同步的视频。
4场景连续性:提供场景截图,生成视觉上匹配的延续内容。
5多模态控制:组合角色图像 + 动态视频 + 背景音频,实现丰富的创作控制。
💰

Pricing & Value

Cost analysis

muapiapp$0.30/sec ($1.50 for 5s, $3.00 for 10s, $4.50 for 15s)

Flat per-second billing with no surcharges. Supports multi-modal references (image + video + audio) in a single request.

Fal.ai$0.3024/sec (high) / $0.2419/sec (basic)

Fal.ai charges $0.3024/sec for high quality and $0.2419/sec for basic. muapiapp is roughly the same on high ($0.30/sec) and 13% cheaper on basic ($0.21/sec).

Replicate$0.3024/sec (high) / $0.2419/sec (basic)

Replicate charges the same as Fal.ai — $0.3024/sec (high), $0.2419/sec (basic). muapiapp is competitive on high quality and 13% cheaper on basic.

** Competitor pricing is estimated based on similar model architectures and usage tiers.

⚙️

Technical Details

Configuration schema

Promptstring

Video description. Use @image1…@image9 to reference images, @video1…@video3 for videos, @audio1…@audio3 for audio. To use a character sheet, reference it with @character:<request_id> (from a completed Seedance 2 Character generation). To use a trained Omni Reference character, reference it with @omni-character:<character_id> where character_id is the value returned by Omni Reference Train Character (e.g. char_1775422630065_4vbana). Both methods can be combined in the same prompt. Multiple characters are supported. Example: '@omni-character:char_1775422630065_4vbana walking through a neon-lit city at night'.

Default Value@image1 is the main character reference. A person walking on the beach at sunset, cinematic lighting
Imagem URLarray

最多 9 个参考Imagem URL(JPEG/PNG/WebP)。第 N 张Imagem对应Prompt中的 @imageN。

Default Valuehttps://d3adwkbyhxyrtq.cloudfront.net/webassets/videomodels/seedance-v2.0-omni-reference.png
参考Vídeo URLarray

最多 3 个参考Vídeo片段 URL(MP4,每个最长 15 秒)。第 N 个Vídeo对应Prompt中的 @videoN。

Default Valueundefined
参考Áudio URLarray

Up to 3 reference audio clip URLs (MP3/WAV, total max 15s). Each Nth audio corresponds to @audioN in the prompt.

Default Valueundefined
ProporçãoEnum (6 options)

SaídaVídeo的Proporção。

Default Value16:9
QualidadeEnum (2 options)

Generation quality. 'high' uses the standard model ($0.30/sec output + $0.09/sec per input video second). 'basic' uses the fast model (~2x speed, $0.21/sec output + $0.063/sec per input video second). Video reference inputs incur an additional 30% surcharge based on their combined duration.

Default Valuehigh
Duração (segundos)(秒)int

Video duration in seconds (4–15).

Default Value5
📖

Implementation Guide

Developer documentation

  1. 将参考图像(JPEG/PNG/WebP)作为 'images_list' 上传,最多 9 张。
  2. 可选上传视频片段(MP4,每个最长 15s)作为 'video_files',最多 3 个视频。
  3. 可选上传音频文件(MP3/WAV)作为 'audio_files',最多 3 个文件,总时长最长 15s。
  4. 编写描述场景的Prompt。使用 @image1…@image9 引用图像,使用 @video1…@video3 引用视频,使用 @audio1…@audio3 引用音频。
  5. 设置时长(4–15s)和Proporção。
  6. 轮询或使用 webhook 获取已完成的视频。

Common Questions

Frequently asked

全能参考与Imagem para Vídeo有什么不同?

Imagem para Vídeo会把你的图像作为实际第一帧并将其制作成动画。全能参考会将图像、视频和音频作为创作引导——模型生成与参考素材在视觉上匹配的新场景,而不会将它们作为实际帧使用。它还支持Imagem para Vídeo不支持的视频和音频参考。

如何在Prompt中引用上传的文件?

使用 @image1、@image2 等,按其在 images_list 中的位置引用图像。使用 @video1、@video2,按其在 video_files 中的位置引用视频。使用 @audio1、@audio2,按其在 audio_files 中的位置引用音频。如果不使用 @ 语法,系统会自动将第一个参考素材作为主要参考。

需要提供所有类型的参考素材吗?

不需要。所有参考数组都是可选的。你可以只提供图像、只提供视频、只提供音频,或任意组合。仅使用文本Prompt也有效。

支持哪些文件格式?

图像:JPEG、PNG 或 WebP(最多 9 个)。视频:仅支持 MP4,每个最长 15 秒(最多 3 个)。音频:MP3、WAV 或其他常见格式,总时长最长 15 秒(最多 3 个文件)。所有 URL 必须可公开访问。

费用如何计算?

费用 =(费率 × Saída时长)+(0.3 × 费率 × Entrada视频总时长)。'high' 质量:Saída $0.30/sec。'basic' 质量:Saída $0.21/sec。如果提供 video_files,则按合计Entrada视频时长的每一秒额外收取 30% 附加费。示例:5s Saída(high)+ 两个 5s Entrada视频 = 5×$0.30 + 10×$0.09 = $1.50 + $0.90 = $2.40。

'high' 和 'basic' 质量有什么区别?

'high' 使用标准模型以获得最佳Saída质量。'basic' 使用快速模型,生成速度约为 2 倍,但质量略有降低——适合快速迭代和预览。未包含 quality 字段的现有请求默认使用 'high'。

支持哪些Proporção?

16:9、9:16、1:1、4:3、3:4 和 21:9。默认值为 16:9。