LTX 2.3 Lipsync: AI Lipsync

LTX-2.3 LipSync generates a realistic talking video by synchronizing mouth movements to an input audio clip. It preserves facial identity, head position, lighting, and natural expressions while producing accurate lip motion, subtle blinking, and stable temporal consistency—powered by the upgraded LTX-2.3 architecture.

📝

Overview

About this model

LTX-2.3 LipSync generates a realistic talking video by synchronizing mouth movements to an input audio clip. It preserves facial identity, head position, lighting, and natural expressions.

1AI Avatars
2Video dubbing
3Dialogue replacement
4Character narration
💰

Pricing & Value

Cost analysis

muapiapp$0.26 avg

Based on 1.3 multiplier

** Competitor pricing is estimated based on similar model architectures and usage tiers.

⚙️

Technical Details

Configuration schema

Promptstring

Optional prompt to guide lipsync generation.

Default ValueAnimate natural lip-sync to the provided audio, add subtle blinking and gentle head motion, maintain the original lighting and facial identity.
Image URLstring

URL of the input portrait image.

Default Valuehttps://d3adwkbyhxyrtq.cloudfront.net/webassets/videomodels/ltx-2.3-lipsync.png
Audio URLstring

URL of the input audio file.

Default Valuehttps://d3adwkbyhxyrtq.cloudfront.net/webassets/videomodels/ltx-2.3-lipsync.wav
ResolutionEnum (3 options)

The resolution of the generated video.

Default Value720p
Seedint

Random seed. -1 for random.

Default Value-1
📖

Implementation Guide

Developer documentation

Provide a clear portrait image and a high-quality audio file. You can also provide a prompt to guide the subtle facial expressions and head motion.

Common Questions

Frequently asked

How long can the audio be?

Currently supports audio clips up to 20 seconds.