The catalog
* FAMILY · alibaba/latentsync

LatentSync Lip Sync for Existing Video

LatentSync is a video-to-video model that regenerates mouth movement in footage you already have so it matches a supplied audio track naturally.

1 variantsalibaba provider
* ABOUT LATENTSYNC

Working with LatentSync

LatentSync takes a video and an audio track and rewrites the lip movement to match. Because the input is video rather than a still, the original performance, camera and background all survive — only the mouth changes.

When to choose LatentSync

Choose LatentSync for dubbing and re-voicing real footage. If you are starting from a single photo instead of a clip, EchoMimic is the audio-driven equivalent; PixVerse Lip Sync is the alternative when you also want built-in TTS voices.

* FREQUENTLY ASKED

About LatentSync

What does LatentSync need as input?
An existing video and the audio you want it to match. It is a video-to-video model, so there is no image-only mode.
Is LatentSync suitable for dubbing into another language?
Yes — that is its main use. Supply the translated audio track and the speaker's mouth is regenerated to fit the new speech.
Does LatentSync change anything besides the mouth?
No. The rest of the frame, including the performance and background, is preserved; the model targets the lip region for natural, high-quality sync.