Lip-sync: what the vendor says moves the mouth
Lip-sync records what a vendor says a mouth is following. Ten of the 17 models put something on the page. Some name the thing that drives the movement, some name the feature and stop, and one argues the step should not exist. As of 2026-09-12.
| Model | What the mouth is said to follow | The published wording behind the cell |
|---|---|---|
| D-ID | Not documented by the vendor | A talk gets created; the mouth is never mentioned |
| Hedra | The supplied audio | The character in the image lip-syncs to the audio provided |
| HeyGen | Named, alongside translation | Video translation is described with cloning and lip-sync |
| Kling AI | Named, with nothing driving it | Lip-sync appears as a capability of the model |
| LTX Studio | Not documented by the vendor | Voices sit on Elements; mouths are not discussed |
| Luma Ray | Not documented by the vendor | Sound is absent from the reference, so mouths are too |
| MiniMax | Whoever is on screen | Lip movement follows the speaker in frame |
| PixVerse | Named as the endpoint's purpose | A speech and lip sync endpoint, with no driver given |
| Runway | Not documented by the vendor | The voices page defines voices, not performances |
| SceneMixer | Argued out of existence | No dubbing pass, so no lip-sync patch afterwards |
| Sora 2 | Named only where it fails | Long, complex speeches are unlikely to sync |
| sync-3 | The audio it is given | Lip movement is matched to the audio on the call |
| Synthesia | The spoken content, framed close | Lip sync is generated from the spoken content |
| Veo | Not documented by the vendor | Audio is documented; mouths are not mentioned |
| Vidu | Not documented by the vendor | Three tracks are described, no mouth among them |
| Wan 3.0 | Not documented by the vendor | Audio defaults on; nothing is said about mouths |
| Wan2.2-S2V | The audio input | The result stays synchronised to the audio input |
Inclusion rule. Models whose documentation says something about how a generated mouth relates to speech. A product that never raises the question is still listed, with the cell recorded as unstated. Order. Alphabetical by model name.
1Three grades of answer, and the grade matters more than the tally
A vendor can name the thing a mouth follows, name the feature and leave the driver out, or never raise the subject. Those three answers look alike in a feature table and behave nothing alike on a shoot. Knowing that movement follows a supplied waveform tells a team it can approve the recording first and expect the picture to obey it. Knowing only that the word appears in the documentation tells a team to budget for a second attempt.
So this column keeps the wording instead of a tick. Where a driver is named it is written out; where only the feature is named the cell says that and no more; where the page is silent, nothing is inferred from the fact that the product plainly does something.
2Products built around a supplied track have the easier sentence
Entries whose whole method is to animate a face to a recording have the easiest sentence to write, because the driver is the input. The mouth follows the file that was handed over, and a vendor can say so plainly without committing to anything about quality. One model card goes further and lets a pose track ride alongside the audio, which keeps the body under direction while the face stays tied to the waveform.
Entries that make sound and picture together have a harder time, because the alignment is a property of the model rather than of an input. Two manage a partial answer: movement follows whoever is in frame, or long speeches are warned not to hold. Both repay close reading, because neither promises that a two-hander will come back usable.
3One entry treats the field as a mistake in the question
The most unusual cell in this column argues there is nothing to synchronise. If the line is performed in the delivery language while the shot renders, no translated track is laid over finished footage, and the repair step everybody else documents has no job to do. Read as a claim about method that is coherent; read as a promise about mouths it cannot be settled from a page.
Recording it in this column rather than dismissing it keeps the comparison even. The cell says what the vendor said, and a reader decides whether removing a stage counts as solving it.
4Silence here is the commonest answer and the least defensible
Seven entries never raise the subject on the page that was read. Several are plainly capable of putting a talking person on screen, so the blank belongs to the document rather than to the product. That still matters, because a producer choosing between tools before any money moves has only the documents, and a document that skips the mouth has skipped the part an audience notices first.
The blank has a pattern to it. Pages written for developers calling an endpoint describe parameters, and there is no parameter for a convincing mouth, so the subject falls outside what the page is for. That explains the gap without excusing it.
5One entry at a time on this column
A cell gets a page of its own where the vendor says something specific in it, or where its silence is unusual among the entries answering the same way. The remaining cells are left in the table above, because a page repeating one short phrase would be worse than a row carrying it.
Entries that name the thing a generated mouth is following:
- Hedra — the supplied audio.
- MiniMax — whoever is on screen.
- sync-3 — the audio it is given.
- Synthesia — the spoken content, framed close.
- Wan2.2-S2V — the audio input.
Entries that name the feature and leave the driver out:
- HeyGen — named, alongside translation.
- Kling AI — named, with nothing driving it.
- PixVerse — named as the endpoint's purpose.
- Sora 2 — named only where it fails.
Entries that argue that nothing needs to follow anything:
- SceneMixer — argued out of existence.
Entries that never raise the subject on the page that was read:
- D-ID — not documented by the vendor.
- Runway — not documented by the vendor.
- Wan 3.0 — not documented by the vendor.
- Lip-syncLip-sync is documentedstated without a mechanism
- Lip-syncThe vendor states there is no separate dubbing step and no lip-sync patch afterwardsstated as unnecessary rather than as a feature
- Audio sourceNative speech with lip-sync tied to the speaker who is on screenspeaker-aware
- Lip-syncLip sync and facial expressions are generated from the spoken contentdriven by the script
- A framing condition attached to lip-syncLip sync performs best when the avatar is framed closer in the scene rather than positioned far from the camera
- Audio sourceAvatar videos are driven by audio, and the character in the image will lip-sync and move to the audio providedhanded to the model rather than made by it
- Audio sourceThe accepted pairs are video with audio, video with text, image with audio and image with textsupplied or read from text
- Dialogue timing against clip lengthA four-second shot is said to accommodate one or two short exchanges and an eight-second clip a few more, with long speeches unlikely to sync
- Lip-syncNot documented by the vendor (as of 2026-09-22)not addressed where voices are defined
- Voice sourceText to speech supports both built-in voices and custom voices created from user-provided sample audiocatalogue or supplied sample
- A second control alongside the audioA pose video argument lets the result follow a pose sequence while staying synchronised to the audio
- LanguagesVideo translation is documented for thirty or more languages, with voice cloning and lip-synccounted on the translation endpoint
6Sources
Each cell is read from the vendor page it links to, checked 2026-09-12. The fields sit side by side on the speech table, and what counts as documented is set out on how read. The other fields: Audio source, Voice source. All of them: the field notes.