OmniHuman 1.5: AI Lipsync Tool

Generate realistic talking head video from portrait image and audio using OmniHuman 1.5.

📝

Overview

About this model

OmniHuman 1.5 is a state-of-the-art lipsync and talking head model that animates a portrait image using an input audio track. It achieves high fidelity, realistic lip-syncing, natural facial expressions, and fluid head movements to create realistic speaking or singing videos.

1Digital Avatars: Create realistic talking avatars for virtual presentations and customer engagement.
2Entertainment: Animate portrait photos and artwork to sing or speak with natural lip sync.
3Social Media: Generate engaging video clips with animated characters matching voiceovers.
đź’°

Pricing & Value

Cost analysis

muapiapp$0.045/sec (720p) / $0.060/sec (1080p)

Dynamic per-second billing based on audio duration (5s to 60s).

Fal.aiNot available

Not available

ReplicateNot available

Not available

* Competitor pricing is estimated based on similar model architectures and usage tiers.

⚙️

Technical Details

Configuration schema

Promptstring

Optional prompt to guide lipsync style.

Default ValueMake her sing confidently into a microphone with natural lip sync
Image URLstring

URL of the input portrait image.

Default Valuehttps://cdn.muapi.ai/assets/omnihuman-1-5.jpg
Audio URLstring

URL of the input audio track.

Default Valuehttps://cdn.muapi.ai/assets/omnihuman-1-5.mp3
Output ResolutionEnum (2 options)

Output video resolution.

Default Value1080
Fast Modeboolean

Enable fast generation mode.

Default Valuefalse
đź“–

Implementation Guide

Developer documentation

How to Use OmniHuman 1.5

  1. Prepare Your Inputs:

    • Image URL: Select a clear, high-quality portrait image. Upload or provide its URL.
    • Audio URL: Provide an audio file containing the speech or song that will drive the avatar's lip-sync.
  2. Configure Parameters:

    • Prompt: Optionally add a prompt to guide the style of the lip sync (e.g., 'Make her sing confidently with natural lip sync').
    • Output Resolution: Choose between 720p or 1080p. The default is 1080p.
    • Fast Mode: Enable fast generation mode (pe_fast_mode) for quicker iterations.
    • Seed: Use a custom seed for reproducible results, or set to -1 for random.
  3. Submit Your Request:

    • Send your prepared inputs to the omnihuman-1-5 endpoint as defined in the technical schema.
  4. Receive and Review the Output:

    • Once processing completes, retrieve the output video URL. Review the generated animation and iterate as needed.
âť“

Common Questions

Frequently asked

What is the maximum duration of generation?

The generation duration is determined by the input audio track length, with a minimum of 5 seconds and a maximum of 60 seconds.

What parameters are required?

Both `image_url` and `audio_url` are mandatory parameters to generate a talking head video.

How is the credit cost calculated?

Cost is calculated dynamically per second of the input audio's duration: $0.045 per second for 720p resolution and $0.06 per second for 1080p resolution.