About this model
Gemini Omni 이미지 투 비디오使用 Google 原生多模态 any-to-any 模型,通过文本提示为一张或多张参考图制作动画。模型会在各帧之间保持主体身份,同时在同一次前向传递中原生生成同步音频、对白、环境声和音乐。它支持 4、6、8 或 10 秒的视频片段,分辨率可选 360p、720p、1080p 或 4K,宽高比可选 16:9 或 9:16,并按출력 결과视频秒数计费。
Cost analysis
| Provider | Cost | Notes |
|---|---|---|
| muapi | $0.039–$0.39 per second of output (by resolution) | Billed per second of output video: $0.039/s at 360p, $0.13/s at 720p, $0.195/s at 1080p, $0.39/s at 4K. Synchronized audio included at no extra charge. |
| Fal.ai | 유사한 초당 요금 | 相同的底层模型(google/gemini-omni-flash/v1.1/image-to-video)。 |
| Replicate | 제공되지 않음 | Gemini Omni Image to Video is not currently available on Replicate. |
Billed per second of output video: $0.039/s at 360p, $0.13/s at 720p, $0.195/s at 1080p, $0.39/s at 4K. Synchronized audio included at no extra charge.
相同的底层模型(google/gemini-omni-flash/v1.1/image-to-video)。
Gemini Omni Image to Video is not currently available on Replicate.
** Competitor pricing is estimated based on similar model architectures and usage tiers.
Configuration schema
| Parameter | Type | Description | Default |
|---|---|---|---|
| 프롬프트 | string | 所需运动和场景的文本描述。Gemini Omni 지원丰富的多模态프롬프트,包括镜头指令、对话和环境오디오提示。 | The suitcase opens by itself and tiny landscapes start unfolding out of it—mountains, forests, oceans, entire cities. Each world expands outward onto the platform, growing larger and larger while miniature weather systems form above them. |
| 参考이미지 | array | Upload 1–7 reference images for the video. Maximum 20 MB each. | https://cdn.muapi.ai/assets/gemini-omni-image-to-video.jpg |
| 재생 시간 (초)(秒) | Enum (4 options) | 生成비디오的재생 시간(秒)。 | 8 |
| 해상도 | Enum (4 options) | 출력비디오해상도。按출력秒数计费:360p 为 $0.039/秒,720p 为 $0.13/秒,1080p 为 $0.195/秒,4K 为 $0.39/秒。 | 1080p |
| 화면 비율 | Enum (2 options) | 출력비디오的화면 비율。 | 16:9 |
| 오디오 ID | array | Up to 3 voice profile IDs returned by the Gemini Omni Audio endpoint. | - |
| 种子 | int | 랜덤 시드(0–2147483647)。固定后可复现结果,但由于模型随机性,结果仍可能有所不同。 | 0 |
| 角色 ID | array | Up to 3 character IDs from Gemini Omni Character to feature in the video. | - |
所需运动和场景的文本描述。Gemini Omni 지원丰富的多模态프롬프트,包括镜头指令、对话和环境오디오提示。
The suitcase opens by itself and tiny landscapes start unfolding out of it—mountains, forests, oceans, entire cities. Each world expands outward onto the platform, growing larger and larger while miniature weather systems form above them.Upload 1–7 reference images for the video. Maximum 20 MB each.
https://cdn.muapi.ai/assets/gemini-omni-image-to-video.jpg生成비디오的재생 시간(秒)。
8출력비디오해상도。按출력秒数计费:360p 为 $0.039/秒,720p 为 $0.13/秒,1080p 为 $0.195/秒,4K 为 $0.39/秒。
1080p출력비디오的화면 비율。
16:9Up to 3 voice profile IDs returned by the Gemini Omni Audio endpoint.
-랜덤 시드(0–2147483647)。固定后可复现结果,但由于模型随机性,结果仍可能有所不同。
0Up to 3 character IDs from Gemini Omni Character to feature in the video.
-Developer documentation
上传参考图
通过 image_urls 提供 1–5 张图像。每张图像都会作为视觉锚点,模型会在各帧之间保持主体身份。
编写动作和场景프롬프트 描述视频中发生的事情——动作、场景、光照和音频提示。例如:'The subject slowly turns to face the camera as golden-hour light sweeps across the scene, leaves rustling in the breeze.'
选择时长和分辨率
选择 4、6、8 或 10 秒。选择 360p 获得快速、低成本的草稿,选择 720p / 1080p 获得标准출력 결과,或选择 4K 获得更高分辨率。
选择宽高比
16:9 — 宽屏、电影感9:16 — 竖屏、移动端优先提交并轮询
POST 到 /api/v1/gemini-omni-image-to-video,并轮询 GET /api/v1/predictions/{request_id}/result,直到 status 为 completed。
请求示例
curl -X POST https://api.muapi.ai/api/v1/gemini-omni-image-to-video \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"prompt": "The subject slowly turns to face the camera as golden-hour light sweeps across the scene.",
"image_urls": ["https://example.com/reference.jpg"],
"duration": 8,
"resolution": "1080p"
}'
Frequently asked
1 到 5 张。每张图像计为 7 单位容量中的 1 个单位(视频使用 2 个单位,角色 ID 使用 1 个单位)。
会——Gemini Omni 이미지 투 비디오旨在让主体在整个生成片段中保持参考图中的身份和外观。
会——同步对白、环境声和音乐会与视频在同一次前向传递中原生生成。
支持 4、6、8 或 10 秒,以及 360p、720p、1080p 或 4K——每种情况都按출력 결과视频秒数计费。
可以——选择 16:9 获得宽屏출력 결과,或选择 9:16 获得竖屏/移动端优先출력 결과。