Sentioscope

Speech and voice controls, as each vendor documents them

Lip-sync: what the vendor says moves the mouth

Lip-sync records what a vendor says a mouth is following. Ten of the 17 models put something on the page. Some name the thing that drives the movement, some name the feature and stop, and one argues the step should not exist. As of 2026-09-12.

Controls documented, model by modelFilled where the model documents that control, hollow where nothing is published about it.Controls documented, model by modelWhat the mouth is said to followWhat the mouth issaid to followThe published wording behind the cellThe publishedwording behind th…D-IDD-ID — What the mouth is said to follow: Not documented by the vendorD-ID — The published wording behind the cell: A talk gets created; the mouth is never mentionedHedraHedra — What the mouth is said to follow: The supplied audioHedra — The published wording behind the cell: The character in the image lip-syncs to the audio providedHeyGenHeyGen — What the mouth is said to follow: Named, alongside translationHeyGen — The published wording behind the cell: Video translation is described with cloning and lip-syncKling AIKling AI — What the mouth is said to follow: Named, with nothing driving itKling AI — The published wording behind the cell: Lip-sync appears as a capability of the modelLTX StudioLTX Studio — What the mouth is said to follow: Not documented by the vendorLTX Studio — The published wording behind the cell: Voices sit on Elements; mouths are not discussedLuma RayLuma Ray — What the mouth is said to follow: Not documented by the vendorLuma Ray — The published wording behind the cell: Sound is absent from the reference, so mouths are tooMiniMaxMiniMax — What the mouth is said to follow: Whoever is on screenMiniMax — The published wording behind the cell: Lip movement follows the speaker in framePixVersePixVerse — What the mouth is said to follow: Named as the endpoint's purposePixVerse — The published wording behind the cell: A speech and lip sync endpoint, with no driver givenRunwayRunway — What the mouth is said to follow: Not documented by the vendorRunway — The published wording behind the cell: The voices page defines voices, not performancesSceneMixerSceneMixer — What the mouth is said to follow: Argued out of existenceSceneMixer — The published wording behind the cell: No dubbing pass, so no lip-sync patch afterwardsSora 2Sora 2 — What the mouth is said to follow: Named only where it failsSora 2 — The published wording behind the cell: Long, complex speeches are unlikely to syncsync-3sync-3 — What the mouth is said to follow: The audio it is givensync-3 — The published wording behind the cell: Lip movement is matched to the audio on the callSynthesiaSynthesia — What the mouth is said to follow: The spoken content, framed closeSynthesia — The published wording behind the cell: Lip sync is generated from the spoken contentVeoVeo — What the mouth is said to follow: Not documented by the vendorVeo — The published wording behind the cell: Audio is documented; mouths are not mentionedViduVidu — What the mouth is said to follow: Not documented by the vendorVidu — The published wording behind the cell: Three tracks are described, no mouth among themWan 3.0Wan 3.0 — What the mouth is said to follow: Not documented by the vendorWan 3.0 — The published wording behind the cell: Audio defaults on; nothing is said about mouthsWan2.2-S2VWan2.2-S2V — What the mouth is said to follow: The audio inputWan2.2-S2V — The published wording behind the cell: The result stays synchronised to the audio input
Fig. 1 Filled where the model documents that control, hollow where nothing is published about it.
Models documenting each controlHow many models document each control. A hollow column is a statement about documentation, not capability.Models documenting each controlWhat the mouth is said to follow6 of 17The published wording behind the ce…The published wording behind the cell16 of 17
Fig. 2 How many models document each control. A hollow column is a statement about documentation, not capability.
What each vendor says a generated mouth is following, if anything. Recorded 2026-09-12.
ModelWhat the mouth is said to followThe published wording behind the cell
D-IDNot documented by the vendorA talk gets created; the mouth is never mentioned
HedraThe supplied audioThe character in the image lip-syncs to the audio provided
HeyGenNamed, alongside translationVideo translation is described with cloning and lip-sync
Kling AINamed, with nothing driving itLip-sync appears as a capability of the model
LTX StudioNot documented by the vendorVoices sit on Elements; mouths are not discussed
Luma RayNot documented by the vendorSound is absent from the reference, so mouths are too
MiniMaxWhoever is on screenLip movement follows the speaker in frame
PixVerseNamed as the endpoint's purposeA speech and lip sync endpoint, with no driver given
RunwayNot documented by the vendorThe voices page defines voices, not performances
SceneMixerArgued out of existenceNo dubbing pass, so no lip-sync patch afterwards
Sora 2Named only where it failsLong, complex speeches are unlikely to sync
sync-3The audio it is givenLip movement is matched to the audio on the call
SynthesiaThe spoken content, framed closeLip sync is generated from the spoken content
VeoNot documented by the vendorAudio is documented; mouths are not mentioned
ViduNot documented by the vendorThree tracks are described, no mouth among them
Wan 3.0Not documented by the vendorAudio defaults on; nothing is said about mouths
Wan2.2-S2VThe audio inputThe result stays synchronised to the audio input

Inclusion rule. Models whose documentation says something about how a generated mouth relates to speech. A product that never raises the question is still listed, with the cell recorded as unstated. Order. Alphabetical by model name.

1Three grades of answer, and the grade matters more than the tally

A vendor can name the thing a mouth follows, name the feature and leave the driver out, or never raise the subject. Those three answers look alike in a feature table and behave nothing alike on a shoot. Knowing that movement follows a supplied waveform tells a team it can approve the recording first and expect the picture to obey it. Knowing only that the word appears in the documentation tells a team to budget for a second attempt.

So this column keeps the wording instead of a tick. Where a driver is named it is written out; where only the feature is named the cell says that and no more; where the page is silent, nothing is inferred from the fact that the product plainly does something.

2Products built around a supplied track have the easier sentence

Entries whose whole method is to animate a face to a recording have the easiest sentence to write, because the driver is the input. The mouth follows the file that was handed over, and a vendor can say so plainly without committing to anything about quality. One model card goes further and lets a pose track ride alongside the audio, which keeps the body under direction while the face stays tied to the waveform.

Entries that make sound and picture together have a harder time, because the alignment is a property of the model rather than of an input. Two manage a partial answer: movement follows whoever is in frame, or long speeches are warned not to hold. Both repay close reading, because neither promises that a two-hander will come back usable.

3One entry treats the field as a mistake in the question

The most unusual cell in this column argues there is nothing to synchronise. If the line is performed in the delivery language while the shot renders, no translated track is laid over finished footage, and the repair step everybody else documents has no job to do. Read as a claim about method that is coherent; read as a promise about mouths it cannot be settled from a page.

Recording it in this column rather than dismissing it keeps the comparison even. The cell says what the vendor said, and a reader decides whether removing a stage counts as solving it.

4Silence here is the commonest answer and the least defensible

Seven entries never raise the subject on the page that was read. Several are plainly capable of putting a talking person on screen, so the blank belongs to the document rather than to the product. That still matters, because a producer choosing between tools before any money moves has only the documents, and a document that skips the mouth has skipped the part an audience notices first.

The blank has a pattern to it. Pages written for developers calling an endpoint describe parameters, and there is no parameter for a convincing mouth, so the subject falls outside what the page is for. That explains the gap without excusing it.

5One entry at a time on this column

A cell gets a page of its own where the vendor says something specific in it, or where its silence is unusual among the entries answering the same way. The remaining cells are left in the table above, because a page repeating one short phrase would be worse than a row carrying it.

Entries that name the thing a generated mouth is following:

Entries that name the feature and leave the driver out:

  • HeyGen — named, alongside translation.
  • Kling AI — named, with nothing driving it.
  • PixVerse — named as the endpoint's purpose.
  • Sora 2 — named only where it fails.

Entries that argue that nothing needs to follow anything:

Entries that never raise the subject on the page that was read:

  • D-ID — not documented by the vendor.
  • Runway — not documented by the vendor.
  • Wan 3.0 — not documented by the vendor.
  • Lip-sync
    Lip-sync is documentedstated without a mechanismKling AI, model guide / recorded 2026-09-12
  • Lip-sync
    The vendor states there is no separate dubbing step and no lip-sync patch afterwardsstated as unnecessary rather than as a featureSceneMixer, languages guide / recorded 2026-09-12
  • Audio source
    Native speech with lip-sync tied to the speaker who is on screenspeaker-awareMiniMax, video generation guide / recorded 2026-09-12
  • Lip-sync
    Lip sync and facial expressions are generated from the spoken contentdriven by the scriptSynthesia, create an avatar / recorded 2026-09-22
  • A framing condition attached to lip-sync
    Lip sync performs best when the avatar is framed closer in the scene rather than positioned far from the cameraSynthesia, create an avatar / recorded 2026-09-22
  • Audio source
    Avatar videos are driven by audio, and the character in the image will lip-sync and move to the audio providedhanded to the model rather than made by itHedra, avatar video guide / recorded 2026-09-22
  • Audio source
    The accepted pairs are video with audio, video with text, image with audio and image with textsupplied or read from textSync, sync-3 model documentation / recorded 2026-09-22
  • Dialogue timing against clip length
    A four-second shot is said to accommodate one or two short exchanges and an eight-second clip a few more, with long speeches unlikely to syncOpenAI, Sora 2 prompting guide / recorded 2026-09-22
  • Lip-sync
    Not documented by the vendor (as of 2026-09-22)not addressed where voices are definedRunway, custom voices reference / recorded 2026-09-22
  • Voice source
    Text to speech supports both built-in voices and custom voices created from user-provided sample audiocatalogue or supplied samplePixVerse, speech and lip sync guide / recorded 2026-09-22
  • A second control alongside the audio
    A pose video argument lets the result follow a pose sequence while staying synchronised to the audioWan-AI, Wan2.2-S2V-14B model card / recorded 2026-09-22
  • Languages
    Video translation is documented for thirty or more languages, with voice cloning and lip-synccounted on the translation endpointHeyGen, API quick start / recorded 2026-09-22

6Sources

Each cell is read from the vendor page it links to, checked 2026-09-12. The fields sit side by side on the speech table, and what counts as documented is set out on how read. The other fields: Audio source, Voice source. All of them: the field notes.