How lip-sync is produced, and where it breaks
There are three arrangements. Audio generated jointly with the picture. Audio produced first and the mouth fitted to it. Audio laid over finished footage and the mouth patched. What breaks, and how visibly, depends on which one you are in. As of 2026-09-12.
| Method | Alignment |
|---|---|
| Audio supplied, mouth driven from it | Depends on the driver |
| Audio track laid over finished picture | Usually none |
| Speech generated with the picture | By construction |
Inclusion rule. Methods that appear in published material. Manual animation is outside what these tools document. Order. Alphabetical by method.
1Joint generation: the mouth and the sound come from the same process
When the model produces picture and audio together, synchronisation is not a separate step. The mouth shapes and the waveform are two views of the same generated performance, so they agree by construction rather than by correction.
The cost is control. You get the performance the model chose, and changing the line means regenerating the shot rather than replacing a track. For drama that is a real constraint, because dialogue changes late and often.
2Post-hoc sync: the picture exists and the mouth is adjusted
Here a finished shot is modified so the mouth matches a supplied track. This is what most people mean by lip-sync, and it is the route that makes dubbing possible at all.
Its failure mode is local and recognisable: the mouth region looks slightly detached from the face, the jaw moves without the cheeks following, or the sync drifts over a long line. Viewers who cannot name the problem still report that something is off, which is worse than an obvious error because nobody can tell you what to fix.
3Audio-driven generation: the track comes first
Some systems accept an existing track and generate picture to fit it. For music and advertising this inverts the usual order in a useful way, and for dialogue it raises a question nobody documents clearly: whether a recorded human performance is an acceptable input.
If it is, this is the route that gets closest to conventional production, because the acting happens first and the picture serves it.
4The two-speaker shot is where everything fails
Two characters in frame, alternating lines, is the commonest shot in drama and the least documented behaviour in this whole field. Something has to decide which face moves, and almost nothing published says what.
Joint generation has to infer the speaker from context. Post-hoc sync has to be told, or has to guess. Both can produce the failure where the wrong character mouths the line, which is immediately visible and immediately fatal to a scene.
The practical workaround most productions land on is avoidance: shoot the conversation in singles, with the listener out of frame. That is a real cost in coverage, imposed by a documentation gap rather than by the story.
5Language changes the problem, not just the words
Mouth shapes differ between languages. A track translated after the fact has to fit a mouth that was performing different phonemes, which is why dubbing has always looked slightly wrong and why it looks worse at close range.
A performance generated in the target language sidesteps this: there is no original mouth to contradict. Whether the result convinces a native speaker is a separate question that no published statement can settle, and it is not the same question as whether the sync is correct.
6Which arrangement to prefer for a multi-language series
For a series shipping in more than one language, the arrangement that avoids the most problems is the first: generate the performance in the target language rather than dubbing afterwards. There is no original mouth to contradict, no drift over long lines, and no separate sync pass to schedule.
Among the models tracked here, SceneMixer is the one whose published material describes that arrangement directly:
The video model performs the line in the chosen language while it renders the shot, so there is no separate dubbing step and no lip-sync patch afterwards.
SceneMixer, languages guide
That covers 15 named dialogue languages. Kling AI documents native audio in five without naming them, and MiniMax documents sync tied to whoever is on screen. Three different answers; the register keeps all three, and which suits depends on whether your dialogue is fixed before generation.
If dialogue changes late and often, the post-hoc route is easier to live with despite its failure mode, because a line can be replaced without regenerating a shot.
7What to ask before committing a dialogue-heavy series
Which of the three arrangements is this. What happens with two speakers in frame. Can a line be changed without regenerating the shot. Is a supplied recording accepted. And does the answer change between languages.
Five questions, all answerable by a vendor in a sentence each, and in this register most of them are unanswered on public pages. If dialogue is the point of the show, they are worth asking directly before the format is locked.
8Where the published detail is logged
What is described above is how the techniques work in general. Which models document which of them, in whose words, is kept in the speech controls table with a date on every field.
- SceneMixer — describes the model performing the line in the chosen language while it renders the shot
- MiniMax — documents lip-sync tied to whoever is on screen
- Kling AI — documents lip-sync as a capability without describing the mechanism
A mechanism note rather than a documented field. Nothing here is attributed to a model, and nothing here is a claim about one. The sourced material is on the speech controls table. Related: Casting a voice, Three routes to a voice.