OmniHuman 1.5: AI Lipsync

Create talking-head videos with Muapi OmniHuman 1.5. Animate a portrait with audio, natural expressions, and lip sync, then retrieve the result by API.

Interactive model controls

Generate realistic talking head video from portrait image and audio using OmniHuman 1.5.

📝

Overview

About this model

OmniHuman 1.5 is a state-of-the-art lipsync and talking head model that animates a portrait image using an input audio track. It achieves high fidelity, realistic lip-syncing, natural facial expressions, and fluid head movements to create realistic speaking or singing videos.

1Digital Avatars: Create realistic talking avatars for virtual presentations and customer engagement.
2Entertainment: Animate portrait photos and artwork to sing or speak with natural lip sync.
3Social Media: Generate engaging video clips with animated characters matching voiceovers.
💰

Pricing & Value

Cost analysis

muapiapp$0.045/sec (720p) / $0.060/sec (1080p)

Dynamic per-second billing based on audio duration (5s to 60s).

Fal.aiNot available

Not available

ReplicateNot available

Not available

** Competitor pricing is estimated based on similar model architectures and usage tiers.

⚙️

Technical Details

Configuration schema

Promptstring

Optional prompt to guide lipsync style.

Default ValueMake her sing confidently into a microphone with natural lip sync
Image URLstring

URL of the input portrait image.

Default Valuehttps://cdn.muapi.ai/assets/omnihuman-1-5.jpg
Audio URLstring

URL of the input audio track.

Default Valuehttps://cdn.muapi.ai/assets/omnihuman-1-5.mp3
Output ResolutionEnum (2 options)

Output video resolution.

Default Value1080
Fast Modeboolean

Enable fast generation mode.

Default Valuefalse
📖

Implementation Guide

Developer documentation

How to Use OmniHuman 1.5

  1. Prepare Your Inputs:

    • Image URL: Select a clear, high-quality portrait image. Upload or provide its URL.
    • Audio URL: Provide an audio file containing the speech or song that will drive the avatar's lip-sync.
  2. Configure Parameters:

    • Prompt: Optionally add a prompt to guide the style of the lip sync (e.g., 'Make her sing confidently with natural lip sync').
    • Output Resolution: Choose between 720p or 1080p. The default is 1080p.
    • Fast Mode: Enable fast generation mode (pe_fast_mode) for quicker iterations.
    • Seed: Use a custom seed for reproducible results, or set to -1 for random.
  3. Submit Your Request:

    • Send your prepared inputs to the omnihuman-1-5 endpoint as defined in the technical schema.
  4. Receive and Review the Output:

    • Once processing completes, retrieve the output video URL. Review the generated animation and iterate as needed.
❓

Common Questions

Frequently asked

What is the maximum duration of generation?

The generation duration is determined by the input audio track length, with a minimum of 5 seconds and a maximum of 60 seconds.

What parameters are required?

Both `image_url` and `audio_url` are mandatory parameters to generate a talking head video.

How is the credit cost calculated?

Cost is calculated dynamically per second of the input audio's duration: $0.045 per second for 720p resolution and $0.06 per second for 1080p resolution.