LTX-2.3 LipSync generates a realistic talking video by synchronizing mouth movements to an input audio clip. It preserves facial identity, head position, lighting, and natural expressions while producing accurate lip motion, subtle blinking, and stable temporal consistency—powered by the upgraded LTX-2.3 architecture.
About this model
LTX-2.3 LipSync generates a realistic talking video by synchronizing mouth movements to an input audio clip. It preserves facial identity, head position, lighting, and natural expressions.
Cost analysis
| Provider | Cost | Notes |
|---|---|---|
| muapiapp | $0.26 avg | Based on 1.3 multiplier |
Based on 1.3 multiplier
** Competitor pricing is estimated based on similar model architectures and usage tiers.
Configuration schema
| Parameter | Type | Description | Default |
|---|---|---|---|
| Prompt | string | Optional prompt to guide lipsync generation. | Animate natural lip-sync to the provided audio, add subtle blinking and gentle head motion, maintain the original lighting and facial identity. |
| Image URL | string | URL of the input portrait image. | https://d3adwkbyhxyrtq.cloudfront.net/webassets/videomodels/ltx-2.3-lipsync.png |
| Audio URL | string | URL of the input audio file. | https://d3adwkbyhxyrtq.cloudfront.net/webassets/videomodels/ltx-2.3-lipsync.wav |
| Resolution | Enum (3 options) | The resolution of the generated video. | 720p |
| Seed | int | Random seed. -1 for random. | -1 |
Optional prompt to guide lipsync generation.
Animate natural lip-sync to the provided audio, add subtle blinking and gentle head motion, maintain the original lighting and facial identity.URL of the input portrait image.
https://d3adwkbyhxyrtq.cloudfront.net/webassets/videomodels/ltx-2.3-lipsync.pngURL of the input audio file.
https://d3adwkbyhxyrtq.cloudfront.net/webassets/videomodels/ltx-2.3-lipsync.wavThe resolution of the generated video.
720pRandom seed. -1 for random.
-1Developer documentation
Provide a clear portrait image and a high-quality audio file. You can also provide a prompt to guide the subtle facial expressions and head motion.
Frequently asked
Currently supports audio clips up to 20 seconds.