使用 Gemini Omni 1.1 Flash 根据文本生成视频。一次调用即可添加参考图像、参考视频、语音和角色 ID,并生成原生音频。免费试用。
关于此模型
Gemini Omni 1.1 Flash 是 Gemini Omni 的文生视频端点:输入提示词,并可选叠加最多 7 张参考图像(辅助条件输入,而非起始关键帧)、一段较短的参考视频片段、最多 3 个来自 Gemini Omni Audio 的语音配置,以及最多 3 个来自 Gemini Omni Character 的角色参考。模型会对你提供的各种输入进行联合推理,并在同一生成流程中返回原生同步对白、环境声和音乐的视频——无需单独的音频处理流程。若要让特定起始图像(以及可选的结束图像)动起来,而不是仅根据提示词生成,请使用 Gemini Omni 1.1 Flash Image to Video。如需编辑现有视频,请参阅 Gemini Omni 1.1 Flash Video Edit;如需按标签寻址的多参考图像合成,请参阅 Gemini Omni 1.1 Flash Reference to Video。
成本分析
| 提供商 | 费用 | 备注 |
|---|---|---|
| muapi | $0.10–$0.30 按输出秒数计费 (by 分辨率) | 按次生成, 无需订阅. Covers text-to-video plus optional image, video-reference, voice, and character inputs. |
| Fal.ai | 暂不可用 | Fal.ai 未列出 Gemini Omni 1.1 Flash。 |
| Replicate | 暂不可用 | Replicate 未列出 Gemini Omni 1.1 Flash。 |
按次生成, 无需订阅. Covers text-to-video plus optional image, video-reference, voice, and character inputs.
Fal.ai 未列出 Gemini Omni 1.1 Flash。
Replicate 未列出 Gemini Omni 1.1 Flash。
** 竞品价格根据相似模型架构和使用层级估算。
配置参数
| 参数 | 类型 | 描述 | 默认值 |
|---|---|---|---|
| 提示词 | string | 对目标视频内容的文本描述——画面、镜头方向、对话和环境音提示。 | A street musician plays a violin on a rainy Paris evening, raindrops tap the cobblestones, a slow melancholic melody, distant café chatter. |
| 参考图像 | array | 最多 7 参考图像 used as auxiliary conditioning (not a starting keyframe). Each counts as 1 quota unit against the shared 7-unit total with video and character_ids. | - |
| 参考视频 | string | 参考视频片段,最大 100MB / 30s。计为 2 个配额单位。 | - |
| 视频裁剪开始时间(秒) | number | Start time, (秒), of the reference video window. | 0 |
| 视频裁剪结束时间(秒) | number | End time, (秒), of the reference video window. Window must not exceed 10 秒。 | 10 |
| 音频 ID | array | 最多 3 voice profile IDs from the Gemini Omni Audio endpoint. Each counts as 1 quota unit. | - |
| 角色 ID | array | 最多 3 character IDs from Gemini Omni Character. Each counts as 1 quota unit. | - |
| 时长(秒) | 枚举(4 个选项) | Duration of the 生成的视频 (秒). Ignored when a reference video is provided — output duration is then determined by the model. | 8 |
| 画面比例 | 枚举(2 个选项) | 输出视频的画面比例。 | 16:9 |
| 分辨率 | 枚举(3 个选项) | 输出分辨率。按输出秒数计费:720p 为 $0.10/秒,1080p 为 $0.15/秒,4K 为 $0.30/秒。 | 720p |
| 种子 | int | 随机种子(0–2147483647)。固定后可复现结果,但由于模型随机性,结果仍可能有所不同。 | 0 |
对目标视频内容的文本描述——画面、镜头方向、对话和环境音提示。
A street musician plays a violin on a rainy Paris evening, raindrops tap the cobblestones, a slow melancholic melody, distant café chatter.最多 7 参考图像 used as auxiliary conditioning (not a starting keyframe). Each counts as 1 quota unit against the shared 7-unit total with video and character_ids.
-参考视频片段,最大 100MB / 30s。计为 2 个配额单位。
-Start time, (秒), of the reference video window.
0End time, (秒), of the reference video window. Window must not exceed 10 秒。
10最多 3 voice profile IDs from the Gemini Omni Audio endpoint. Each counts as 1 quota unit.
-最多 3 character IDs from Gemini Omni Character. Each counts as 1 quota unit.
-Duration of the 生成的视频 (秒). Ignored when a reference video is provided — output duration is then determined by the model.
8输出视频的画面比例。
16:9输出分辨率。按输出秒数计费:720p 为 $0.10/秒,1080p 为 $0.15/秒,4K 为 $0.30/秒。
720p随机种子(0–2147483647)。固定后可复现结果,但由于模型随机性,结果仍可能有所不同。
0开发者文档
编写多模态提示词 将视觉内容、镜头指示、对白和环境声一起描述。示例:'巴黎雨夜,一位街头音乐家拉着小提琴,雨滴敲打着鹅卵石,缓慢而忧郁的旋律,远处传来咖啡馆的交谈声。'
叠加可选条件输入
image_urls,最多 7 张),作为辅助视觉条件输入。video_url + trim_start/trim_end,最长 10 秒窗口),用于以连续性为导向的延展。audio_ids)和/或角色参考(character_ids)。遵守共享配额 每张图像计 1 个单位,参考视频计 2 个单位,每个角色 ID 计 1 个单位——总数不得超过 7。
选择时长、分辨率和宽高比 提供参考视频时,时长(4/6/8/10s)会被忽略——模型决定输出长度。分辨率按输出每秒计费:720p 为 $0.10/s,1080p 为 $0.15/s,4K 为 $0.30/s。
提交并轮询
curl -X POST https://api.muapi.ai/api/v1/gemini-omni-flash-1-1-text-to-video \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"prompt": "A street musician plays a violin on a rainy Paris evening",
"image_urls": ["https://example.com/scene.jpg"],
"duration": 8,
"resolution": "1080p",
"aspect_ratio": "16:9"
}'
然后调用 GET /api/v1/predictions/{request_id}/result 轮询,直到 status 为 completed。
提示词技巧
常见问答
这是 Gemini Omni 的文生视频端点:一次调用即可接收提示词、可选参考图像、参考视频片段、语音配置和角色 ID,并返回带原生同步音频的视频。
可以。参考图像、参考视频、语音配置和角色 ID 都可以组合使用,但必须遵守 7 个单位的共享配额。
每张参考图像消耗 1 个单位,参考视频消耗 2 个单位,每个角色 ID 消耗 1 个单位。图像 +(视频 × 2)+ 角色 ID 的总数不得超过 7。
它会被忽略。提供参考视频时,模型会自动决定输出时长,而不是使用请求中的 `duration` 值。
根据输出视频的分辨率按秒计费:720p 为 $0.10/s,1080p 为 $0.15/s,4K 为 $0.30/s。10 秒的 1080p 视频费用为 $1.50。
使用 [Gemini Omni Audio](/playground/gemini-omni-audio) 生成可复用的语音配置,使用 [Gemini Omni Character](/playground/gemini-omni-character) 生成可复用的角色参考,然后将返回的 ID 传入此处。
不应该。请改用 [Gemini Omni 1.1 Flash Image to Video](/playground/gemini-omni-flash-1-1-image-to-video)——其 `first_frame_url` 是实际的起始关键帧。本端点的 `image_urls` 是辅助参考图像,而不是关键帧。