Lip-sync: matching a mouth to a waveform
Lip-sync is the agreement between a mouth on screen and the sound playing over it. Three mechanisms produce it: generating both together, animating a face to supplied audio, or repairing the mouth afterwards. They fail visibly differently. As of 2026-09-12.
| Mechanism | What drives the mouth | How it tends to fail |
|---|---|---|
| Animated to audio | A supplied waveform | Tight timing on a face that does not sit in the shot |
| Generated together | Whatever the model decides | Approximate movement, worst on long speeches |
| Repaired afterwards | A new track over old frames | A mouth that reads as pasted on |
Inclusion rule. The three mechanisms a vendor could be describing when it names lip-sync. A product is not assigned to a mechanism unless its own documentation names the driver. Order. Alphabetical by mechanism.
1Naming the word is not naming the mechanism
Ten entries in this register mention lip-sync and five name something as the driver. Four put the word on the page and stop, and one argues the step away, which leaves a production unable to predict which artefact to expect when a shot comes back wrong.
That difference matters more than it sounds. Knowing the mouth follows a supplied track means a bad result from a good recording is a picture problem, and the diagnosis is over before it starts.
2The two-speaker shot breaks all three differently
Generated together, a model has to decide which of two faces is talking. Animated to audio, it has one waveform and no way to know whose it is. Repaired afterwards, both mouths are candidates and only one should move.
Almost nothing in this register addresses the case. One entry ties movement to whoever is on screen, one asks for speakers to be labelled in the prompt, and one warns that closer framing performs better, which is a quiet way of saying the wide two-shot is the weak case.
3Removing the step is a fourth position
One vendor argues there is nothing to synchronise: if the line is performed in the delivery language while the shot renders, no translated track is laid over finished footage and the repair stage has no job. Read as a claim about method that is coherent.
Read as a promise about mouths it settles nothing, because speech generated with a picture can still be out of step with it. The absence of a repair stage is not evidence that no repair is needed.
4Sources read for this entry
This page defines the term and logs what each vendor documents about it, each figure read on 2026-09-12. Related: Voice id, Voice cloning.