Generate realistic talking head video from portrait image and audio using OmniHuman 1.5.
About this model
OmniHuman 1.5 is a state-of-the-art lipsync and talking head model that animates a portrait image using an input audio track. It achieves high fidelity, realistic lip-syncing, natural facial expressions, and fluid head movements to create realistic speaking or singing videos.
Cost analysis
| Provider | Cost | Notes |
|---|---|---|
| muapiapp | $0.045/sec (720p) / $0.060/sec (1080p) | Dynamic per-second billing based on audio duration (5s to 60s). |
| Fal.ai | Not available | Not available |
| Replicate | Not available | Not available |
Dynamic per-second billing based on audio duration (5s to 60s).
Not available
Not available
* Competitor pricing is estimated based on similar model architectures and usage tiers.
Configuration schema
| Parameter | Type | Description | Default |
|---|---|---|---|
| Prompt | string | Optional prompt to guide lipsync style. | Make her sing confidently into a microphone with natural lip sync |
| Image URL | string | URL of the input portrait image. | https://cdn.muapi.ai/assets/omnihuman-1-5.jpg |
| Audio URL | string | URL of the input audio track. | https://cdn.muapi.ai/assets/omnihuman-1-5.mp3 |
| Output Resolution | Enum (2 options) | Output video resolution. | 1080 |
| Fast Mode | boolean | Enable fast generation mode. | false |
Optional prompt to guide lipsync style.
Make her sing confidently into a microphone with natural lip syncURL of the input portrait image.
https://cdn.muapi.ai/assets/omnihuman-1-5.jpgURL of the input audio track.
https://cdn.muapi.ai/assets/omnihuman-1-5.mp3Output video resolution.
1080Enable fast generation mode.
falseDeveloper documentation
Prepare Your Inputs:
Configure Parameters:
720p or 1080p. The default is 1080p.pe_fast_mode) for quicker iterations.-1 for random.Submit Your Request:
omnihuman-1-5 endpoint as defined in the technical schema.Receive and Review the Output:
Frequently asked
The generation duration is determined by the input audio track length, with a minimum of 5 seconds and a maximum of 60 seconds.
Both `image_url` and `audio_url` are mandatory parameters to generate a talking head video.
Cost is calculated dynamically per second of the input audio's duration: $0.045 per second for 720p resolution and $0.06 per second for 1080p resolution.