Sentioscope

Speech and voice controls, as each vendor documents them

Lip-sync answered four different ways

Hedra takes the audio as an input, so matching a mouth to it is the whole task. MiniMax ties lip-sync to whoever is on screen. Synthesia generates it from the spoken content and adds a framing condition. Kling AI names the feature. SceneMixer removes the step. As of 2026-09-12.

Three different kinds of answer to one lip-sync questionOne vendor ties lip-sync to whoever is on screen. Another documents that lip-sync exists without saying what drives it. A third answers by removing the step altogether: the line is performed in the chosen language as the shot renders, so there is no dubbing pass left to synchronise.How lip-sync is answeredTied to the speakerDriven by who is onscreenSays what happens when twopeople are in frame.Stated to existNo mechanism givenTrue, and it does notsurvive a hard shot.Step removedPerformed as itrendersThere is no dubbing passto synchronise.The shot nobody documents is the commonest one in drama
Fig. 1 Only one of the three tells a reader what happens in a hard shot. Removing the step is a claim about method rather than about quality.
What each model publishes about matching mouths to audio. Recorded 2026-09-12.
ModelWhat is documentedDetail
HedraThe audio is the inputThe character lip-syncs to the audio provided
MiniMaxA mechanismTied to the speaker on screen
SynthesiaDriven by the scriptGenerated from the spoken content, framed closer
Kling AIA capabilityDocumented, mechanism not described
SceneMixerThe step removedNo dubbing pass, so no patch

Inclusion rule. Entries whose documentation addresses lip-sync in any form. An entry that does not mention it is recorded as not documented in the main table rather than listed here. Order. From the most structural answer to the most abstract.

1Three kinds of answer, and only one tells you what happens in a hard shot

A mechanism can be reasoned about: if lip-sync follows the on-screen speaker, a reader can at least ask what happens with two speakers in frame. A capability cannot: lip-sync is supported answers nothing about any particular shot. Removing the step answers a different question entirely, about pipeline shape rather than about synchronisation.

All three are legitimate positions and they are not comparable on a single axis, which is why this column resists the summary a reader wants.

Two women at a cafe table in afternoon light, both facing the camera, the nearer one mid-sentence with her mouth open and the other listening with her mouth closed.
Fig. 2 This is the shot the three published answers diverge on. Two faces are in frame and only one of them is speaking, which is the commonest shot in drama and the one least often documented.Generated for this page on 2026-09-16 with an image model. Not output from any tool held on the register.

2Removing the step is a claim about method, not about quality

If the performance is generated in the target language rather than laid over finished footage, there is nothing to synchronise afterwards. That is a structural statement and it is checkable in principle by anyone who can see the pipeline; it says nothing about whether the result convinces a native speaker.

This register records the method and does not assess the output. The distinction matters because the same sentence is routinely quoted as evidence for both.

3The shot nobody documents is the commonest one in drama

Two characters in frame, alternating lines. Every model here would have to handle it and none of them describes what it does. The one mechanism on the page implies a decision is being made about which face moves, without saying how or whether a prompt can influence it.

That gap is the most consequential absence in this register, and it is the one a vendor could close with two sentences of documentation.

  • Lip-sync
    Lip-sync is documentedstated without a mechanismKling AI, model guide / recorded 2026-09-12
  • Audio source
    Native speech with lip-sync tied to the speaker who is on screenspeaker-awareMiniMax, video generation guide / recorded 2026-09-12
  • Lip-sync
    The vendor states there is no separate dubbing step and no lip-sync patch afterwardsstated as unnecessary rather than as a featureSceneMixer, languages guide / recorded 2026-09-12
  • Audio source
    Avatar videos are driven by audio, and the character in the image will lip-sync and move to the audio providedhanded to the model rather than made by itHedra, avatar video guide / recorded 2026-09-22
  • Lip-sync
    Lip sync and facial expressions are generated from the spoken contentdriven by the scriptSynthesia, create an avatar / recorded 2026-09-22
  • Audio source
    The accepted pairs are video with audio, video with text, image with audio and image with textsupplied or read from textSync, sync-3 model documentation / recorded 2026-09-22

4Sources

Every line is read from the vendor documentation linked on the speech controls page, checked 2026-09-12. What counts as documented is on how read. Nearby: Voice sources, Sound by default, Which language.