Sentioscope

Speech and voice controls, as each vendor documents them

Entries that name what a generated mouth follows

Five entries name a driver. Three of them are handed the waveform, which makes it easy to say. One drives the face from a script and attaches a framing condition. One makes the sound itself and still names what the mouth follows. As of 2026-09-22.

A named driver says which half to fixWhen a shot comes back wrong the useful question is whether the audio or the picture caused it. If the mouth follows a supplied track, a bad result from a good track is a picture problem and the diagnosis is over.A dialogue shot comes back looking wrongThe driver is namedDiagnosis is shortCheck the input the vendor named, thenblame the other halfOnly the feature is namedRegenerate both halvesThree possible mechanisms, and no way totell which appliedFour of these five are audio-driven, which makes the naming cheap
Fig. 1 Without a driver the same shot is a coin toss between three mechanisms that fail differently.
Entries with a named driver, beside the architecture that makes the naming possible. Recorded 2026-09-22.
ModelWhat the mouth followsAnd where the sound comes from
HedraThe supplied audioSupplied; audio is required
MiniMaxWhoever is on screenWith the picture
sync-3The audio it is givenSupplied, or read from text
SynthesiaThe spoken content, framed closeFrom the script, or uploaded
Wan2.2-S2VThe audio inputSupplied; the model is audio-driven

Inclusion rule. Entries whose documentation names the thing a generated mouth is following. A page that names lip-sync as a feature without a driver is on a separate route page. Order. Alphabetical by model name.

1Naming the driver is cheap for some architectures and hard for others

Where audio is an input, saying the mouth follows it describes the interface. Three of these five are in that position, which is why this group leans towards audio-driven products rather than the ones that make sound and picture together.

The two exceptions are instructive. One drives the face from the script, which is a different bargain entirely. One generates speech natively and still names a driver: lip movement follows whoever is on screen, which is a claim about a decision inside the model and the one sentence here touching the commonest shot in drama.

2A driver tells a production which half to fix

When a shot comes back wrong, the useful question is whether the audio or the picture caused it. A named driver answers that before any diagnosis: if the mouth follows the supplied track, then a bad result with a good track is a picture problem.

Without a driver the same shot is a coin toss between three mechanisms that fail differently, and a production ends up regenerating both halves because it cannot tell which one to blame.

3The one script-driven entry publishes the most useful sentence here

Driving lip movement from the spoken content makes a rewrite the only lever, which is a real limitation for drama. The same vendor then says performance improves with the avatar framed closer to camera, which names the weak case.

Naming a weak case is rarer than naming a capability and worth more. Wide framing is the staple of dialogue coverage, so that one sentence constrains a shot list in a way no capability claim does.

4The entries on this route, one page each

Each of these links to that entry on the field this route groups by. The wording behind every cell, and the date it was read, sits on the page it links to.

  • Audio source
    Avatar videos are driven by audio, and the character in the image will lip-sync and move to the audio providedhanded to the model rather than made by itHedra, avatar video guide / recorded 2026-09-22
  • Audio source
    Native speech with lip-sync tied to the speaker who is on screenspeaker-awareMiniMax, video generation guide / recorded 2026-09-12
  • Lip-sync
    Lip sync and facial expressions are generated from the spoken contentdriven by the scriptSynthesia, create an avatar / recorded 2026-09-22
  • A framing condition attached to lip-sync
    Lip sync performs best when the avatar is framed closer in the scene rather than positioned far from the cameraSynthesia, create an avatar / recorded 2026-09-22
  • Audio source
    The accepted pairs are video with audio, video with text, image with audio and image with textsupplied or read from textSync, sync-3 model documentation / recorded 2026-09-22
  • A second control alongside the audio
    A pose video argument lets the result follow a pose sequence while staying synchronised to the audioWan-AI, Wan2.2-S2V-14B model card / recorded 2026-09-22

5Sources

Membership of this route is decided by the wording in the field it groups by, read from the documentation each vendor publishes. The column itself is on the field note, all five columns are on the speech table, and what counts as documented is on how read. Other routes: Named, not explained, The subject skipped.