Animate images into video with Gemini Omni — native synchronized audio, dialogue, and music. Try free — pay per generation, no subscription.
About this model
Gemini Omni Image to Video animates one or more reference images with a text prompt using Google's natively multimodal any-to-any model. Subject identity is preserved across frames while synchronized audio — dialogue, ambient sound, and music — is generated natively in the same forward pass. It supports 4, 6, 8, or 10 second clips at 360p, 720p, 1080p, or 4K, in either 16:9 or 9:16 aspect ratio, and is billed per second of output video.
Cost analysis
| Provider | Cost | Notes |
|---|---|---|
| muapi | $0.039–$0.39 per second of output (by resolution) | Billed per second of output video: $0.039/s at 360p, $0.13/s at 720p, $0.195/s at 1080p, $0.39/s at 4K. Synchronized audio included at no extra charge. |
| Fal.ai | Comparable per-second pricing | Same underlying model (google/gemini-omni-flash/v1.1/image-to-video). |
| Replicate | Not available | Gemini Omni Image to Video is not currently available on Replicate. |
Billed per second of output video: $0.039/s at 360p, $0.13/s at 720p, $0.195/s at 1080p, $0.39/s at 4K. Synchronized audio included at no extra charge.
Same underlying model (google/gemini-omni-flash/v1.1/image-to-video).
Gemini Omni Image to Video is not currently available on Replicate.
** Competitor pricing is estimated based on similar model architectures and usage tiers.
Configuration schema
| Parameter | Type | Description | Default |
|---|---|---|---|
| Prompt | string | Text description of the desired motion and scene. Gemini Omni supports rich multimodal prompts including camera direction, dialogue, and ambient audio cues. | The suitcase opens by itself and tiny landscapes start unfolding out of it—mountains, forests, oceans, entire cities. Each world expands outward onto the platform, growing larger and larger while miniature weather systems form above them. |
| Reference Images | array | Upload 1–7 reference images for the video. Maximum 20 MB each. | https://cdn.muapi.ai/assets/gemini-omni-image-to-video.jpg |
| Duration (seconds) | Enum (4 options) | Duration of the generated video in seconds. | 8 |
| Resolution | Enum (4 options) | Output video resolution. Billed per second of output: $0.039/s at 360p, $0.13/s at 720p, $0.195/s at 1080p, $0.39/s at 4K. | 1080p |
| Aspect Ratio | Enum (2 options) | Output video aspect ratio. | 16:9 |
| Audio IDs | array | Up to 3 voice profile IDs returned by the Gemini Omni Audio endpoint. | - |
| Seed | int | Random seed (0–2147483647). Fix for reproducibility; results may still vary due to model stochasticity. | 0 |
| Character IDs | array | Up to 3 character IDs from Gemini Omni Character to feature in the video. | - |
Text description of the desired motion and scene. Gemini Omni supports rich multimodal prompts including camera direction, dialogue, and ambient audio cues.
The suitcase opens by itself and tiny landscapes start unfolding out of it—mountains, forests, oceans, entire cities. Each world expands outward onto the platform, growing larger and larger while miniature weather systems form above them.Upload 1–7 reference images for the video. Maximum 20 MB each.
https://cdn.muapi.ai/assets/gemini-omni-image-to-video.jpgDuration of the generated video in seconds.
8Output video resolution. Billed per second of output: $0.039/s at 360p, $0.13/s at 720p, $0.195/s at 1080p, $0.39/s at 4K.
1080pOutput video aspect ratio.
16:9Up to 3 voice profile IDs returned by the Gemini Omni Audio endpoint.
-Random seed (0–2147483647). Fix for reproducibility; results may still vary due to model stochasticity.
0Up to 3 character IDs from Gemini Omni Character to feature in the video.
-Developer documentation
Upload reference images
Provide 1–5 images via image_urls. Each image acts as a visual anchor. The model preserves subject identity across frames.
Write a motion and scene prompt Describe what happens in the video — motion, setting, lighting, and audio cues. Example: 'The subject slowly turns to face the camera as golden-hour light sweeps across the scene, leaves rustling in the breeze.'
Choose duration and resolution
Pick 4, 6, 8, or 10 seconds. Choose 360p for a fast, low-cost draft, 720p / 1080p for standard output, or 4K for higher resolution.
Pick an aspect ratio
16:9 — widescreen, cinematic9:16 — vertical, mobile-firstSubmit and poll
POST to /api/v1/gemini-omni-image-to-video and poll GET /api/v1/predictions/{request_id}/result until status is completed.
Example request
curl -X POST https://api.muapi.ai/api/v1/gemini-omni-image-to-video \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"prompt": "The subject slowly turns to face the camera as golden-hour light sweeps across the scene.",
"image_urls": ["https://example.com/reference.jpg"],
"duration": 8,
"resolution": "1080p"
}'
Frequently asked
Between 1 and 5 images. Each image counts as 1 unit toward the 7-unit capacity (videos use 2, character IDs use 1).
Yes — Gemini Omni Image to Video is designed to maintain subject identity and appearance from the reference images throughout the generated clip.
Yes — synchronized dialogue, ambient sound, and music are generated natively alongside the video in the same forward pass.
4, 6, 8, or 10 seconds, at 360p, 720p, 1080p, or 4K — each billed per second of output video.
Yes — choose 16:9 for widescreen or 9:16 for vertical/mobile-first output.