SD 2.0 全能参考可使用参考图像、视频和音频生成具有视觉一致性的视频。保持角色身份、视觉风格和场景连续性。单次请求最多支持 9 张图像、3 个视频片段和 3 个音频文件。可在提示词中使用 @image1、@video1、@audio1 语法,精确控制每个参考素材对生成视频的影响。
关于此模型
SD 2.0 全能参考可使用参考图像、视频和音频生成具有视觉一致性的视频。不同于只将单张图像制作成动画的标准图生视频,全能参考会将你上传的素材用作创作引导——在保持角色身份、视觉风格和场景连续性的同时生成新视频。单次请求可组合最多 9 张图像、3 个视频片段和 3 个音频文件。在提示词中使用 @image1、@video1、@audio1 语法,可精确控制每个参考素材如何影响生成的视频。
成本分析
| 提供商 | 费用 | 备注 |
|---|---|---|
| muapiapp | $0.30/秒 ($1.50 每 5 秒, $3.00 for 10s, $4.50 for 15s) | Flat per-second billing with no sur收费. Supports multi-modal references (image + video + audio) in a single request. |
| Fal.ai | $0.3024/秒 (高质量) / $0.2419/秒 (基础) | Fal.ai 收费 $0.3024/sec for high 质量 and $0.2419/sec for basic. muapiapp is roughly the same on high ($0.30/sec) and 13% 更便宜 on basic ($0.21/sec). |
| Replicate | $0.3024/秒 (高质量) / $0.2419/秒 (基础) | Replicate 收费 the same as Fal.ai — $0.3024/sec (high), $0.2419/sec (basic). muapiapp is competitive on high 质量 and 13% 更便宜 on basic. |
Flat per-second billing with no sur收费. Supports multi-modal references (image + video + audio) in a single request.
Fal.ai 收费 $0.3024/sec for high 质量 and $0.2419/sec for basic. muapiapp is roughly the same on high ($0.30/sec) and 13% 更便宜 on basic ($0.21/sec).
Replicate 收费 the same as Fal.ai — $0.3024/sec (high), $0.2419/sec (basic). muapiapp is competitive on high 质量 and 13% 更便宜 on basic.
** 竞品价格根据相似模型架构和使用层级估算。
配置参数
| 参数 | 类型 | 描述 | 默认值 |
|---|---|---|---|
| 提示词 | string | Video description. Use @image1…@image9 to 参考图像, @video1…@video3 for videos, @audio1…@audio3 for audio. To use a character sheet, reference it with @character:<request_id> (from a completed Seedance 2 Character generation). To use a trained Omni Reference character, reference it with @omni-character:<character_id> where character_id is the value returned by Omni Reference Train Character (e.g. char_1775422630065_4vbana). Both methods can be combined in the same prompt. Multiple characters are supported. Example: '@omni-character:char_1775422630065_4vbana walking through a neon-lit city at night'. | @image1 is the main character reference. A person walking on the beach at sunset, cinematic lighting |
| 图像 URL | array | 最多 9 个参考图像 URL(JPEG/PNG/WebP)。第 N 张图像对应提示词中的 @imageN。 | https://d3adwkbyhxyrtq.cloudfront.net/webassets/videomodels/seedance-v2.0-omni-reference.png |
| 参考视频 URL | array | 最多 3 个参考视频片段 URL(MP4,每个最长 15 秒)。第 N 个视频对应提示词中的 @videoN。 | undefined |
| 参考音频 URL | array | 最多 3 参考音频 clip URLs (MP3/WAV, total max 15s). Each Nth audio corresponds to @audioN in the prompt. | undefined |
| 画面比例 | 枚举(6 个选项) | 输出视频的画面比例。 | 16:9 |
| 质量 | 枚举(2 个选项) | Generation quality. 'high' uses the standard model ($0.30/sec output + $0.09/sec per 输入视频 second). 'basic' uses the fast model (~2x speed, $0.21/sec output + $0.063/sec per 输入视频 second). Video reference inputs incur an additional 30% surcharge based on their combined duration. | high |
| 时长(秒) | int | Video duration (秒) (4–15). | 5 |
Video description. Use @image1…@image9 to 参考图像, @video1…@video3 for videos, @audio1…@audio3 for audio. To use a character sheet, reference it with @character:<request_id> (from a completed Seedance 2 Character generation). To use a trained Omni Reference character, reference it with @omni-character:<character_id> where character_id is the value returned by Omni Reference Train Character (e.g. char_1775422630065_4vbana). Both methods can be combined in the same prompt. Multiple characters are supported. Example: '@omni-character:char_1775422630065_4vbana walking through a neon-lit city at night'.
@image1 is the main character reference. A person walking on the beach at sunset, cinematic lighting最多 9 个参考图像 URL(JPEG/PNG/WebP)。第 N 张图像对应提示词中的 @imageN。
https://d3adwkbyhxyrtq.cloudfront.net/webassets/videomodels/seedance-v2.0-omni-reference.png最多 3 个参考视频片段 URL(MP4,每个最长 15 秒)。第 N 个视频对应提示词中的 @videoN。
undefined最多 3 参考音频 clip URLs (MP3/WAV, total max 15s). Each Nth audio corresponds to @audioN in the prompt.
undefined输出视频的画面比例。
16:9Generation quality. 'high' uses the standard model ($0.30/sec output + $0.09/sec per 输入视频 second). 'basic' uses the fast model (~2x speed, $0.21/sec output + $0.063/sec per 输入视频 second). Video reference inputs incur an additional 30% surcharge based on their combined duration.
highVideo duration (秒) (4–15).
5开发者文档
常见问答
图生视频会把你的图像作为实际第一帧并将其制作成动画。全能参考会将图像、视频和音频作为创作引导——模型生成与参考素材在视觉上匹配的新场景,而不会将它们作为实际帧使用。它还支持图生视频不支持的视频和音频参考。
使用 @image1、@image2 等,按其在 images_list 中的位置引用图像。使用 @video1、@video2,按其在 video_files 中的位置引用视频。使用 @audio1、@audio2,按其在 audio_files 中的位置引用音频。如果不使用 @ 语法,系统会自动将第一个参考素材作为主要参考。
不需要。所有参考数组都是可选的。你可以只提供图像、只提供视频、只提供音频,或任意组合。仅使用文本提示词也有效。
图像:JPEG、PNG 或 WebP(最多 9 个)。视频:仅支持 MP4,每个最长 15 秒(最多 3 个)。音频:MP3、WAV 或其他常见格式,总时长最长 15 秒(最多 3 个文件)。所有 URL 必须可公开访问。
费用 =(费率 × 输出时长)+(0.3 × 费率 × 输入视频总时长)。'high' 质量:输出 $0.30/sec。'basic' 质量:输出 $0.21/sec。如果提供 video_files,则按合计输入视频时长的每一秒额外收取 30% 附加费。示例:5s 输出(high)+ 两个 5s 输入视频 = 5×$0.30 + 10×$0.09 = $1.50 + $0.90 = $2.40。
'high' 使用标准模型以获得最佳输出质量。'basic' 使用快速模型,生成速度约为 2 倍,但质量略有降低——适合快速迭代和预览。未包含 quality 字段的现有请求默认使用 'high'。
16:9、9:16、1:1、4:3、3:4 和 21:9。默认值为 16:9。