Lip-sync answered four different ways
Hedra takes the audio as an input, so matching a mouth to it is the whole task. MiniMax ties lip-sync to whoever is on screen. Synthesia generates it from the spoken content and adds a framing condition. Kling AI names the feature. SceneMixer removes the step. As of 2026-09-12.
| Model | What is documented | Detail |
|---|---|---|
| Hedra | The audio is the input | The character lip-syncs to the audio provided |
| MiniMax | A mechanism | Tied to the speaker on screen |
| Synthesia | Driven by the script | Generated from the spoken content, framed closer |
| Kling AI | A capability | Documented, mechanism not described |
| SceneMixer | The step removed | No dubbing pass, so no patch |
Inclusion rule. Entries whose documentation addresses lip-sync in any form. An entry that does not mention it is recorded as not documented in the main table rather than listed here. Order. From the most structural answer to the most abstract.
1Three kinds of answer, and only one tells you what happens in a hard shot
A mechanism can be reasoned about: if lip-sync follows the on-screen speaker, a reader can at least ask what happens with two speakers in frame. A capability cannot: lip-sync is supported answers nothing about any particular shot. Removing the step answers a different question entirely, about pipeline shape rather than about synchronisation.
All three are legitimate positions and they are not comparable on a single axis, which is why this column resists the summary a reader wants.

2Removing the step is a claim about method, not about quality
If the performance is generated in the target language rather than laid over finished footage, there is nothing to synchronise afterwards. That is a structural statement and it is checkable in principle by anyone who can see the pipeline; it says nothing about whether the result convinces a native speaker.
This register records the method and does not assess the output. The distinction matters because the same sentence is routinely quoted as evidence for both.
3The shot nobody documents is the commonest one in drama
Two characters in frame, alternating lines. Every model here would have to handle it and none of them describes what it does. The one mechanism on the page implies a decision is being made about which face moves, without saying how or whether a prompt can influence it.
That gap is the most consequential absence in this register, and it is the one a vendor could close with two sentences of documentation.
- Lip-syncLip-sync is documentedstated without a mechanism
- Audio sourceNative speech with lip-sync tied to the speaker who is on screenspeaker-aware
- Lip-syncThe vendor states there is no separate dubbing step and no lip-sync patch afterwardsstated as unnecessary rather than as a feature
- Audio sourceAvatar videos are driven by audio, and the character in the image will lip-sync and move to the audio providedhanded to the model rather than made by it
- Lip-syncLip sync and facial expressions are generated from the spoken contentdriven by the script
- Audio sourceThe accepted pairs are video with audio, video with text, image with audio and image with textsupplied or read from text
4Sources
Every line is read from the vendor documentation linked on the speech controls page, checked 2026-09-12. What counts as documented is on how read. Nearby: Voice sources, Sound by default, Which language.