Entries that name what a generated mouth follows
Five entries name a driver. Three of them are handed the waveform, which makes it easy to say. One drives the face from a script and attaches a framing condition. One makes the sound itself and still names what the mouth follows. As of 2026-09-22.
| Model | What the mouth follows | And where the sound comes from |
|---|---|---|
| Hedra | The supplied audio | Supplied; audio is required |
| MiniMax | Whoever is on screen | With the picture |
| sync-3 | The audio it is given | Supplied, or read from text |
| Synthesia | The spoken content, framed close | From the script, or uploaded |
| Wan2.2-S2V | The audio input | Supplied; the model is audio-driven |
Inclusion rule. Entries whose documentation names the thing a generated mouth is following. A page that names lip-sync as a feature without a driver is on a separate route page. Order. Alphabetical by model name.
1Naming the driver is cheap for some architectures and hard for others
Where audio is an input, saying the mouth follows it describes the interface. Three of these five are in that position, which is why this group leans towards audio-driven products rather than the ones that make sound and picture together.
The two exceptions are instructive. One drives the face from the script, which is a different bargain entirely. One generates speech natively and still names a driver: lip movement follows whoever is on screen, which is a claim about a decision inside the model and the one sentence here touching the commonest shot in drama.
2A driver tells a production which half to fix
When a shot comes back wrong, the useful question is whether the audio or the picture caused it. A named driver answers that before any diagnosis: if the mouth follows the supplied track, then a bad result with a good track is a picture problem.
Without a driver the same shot is a coin toss between three mechanisms that fail differently, and a production ends up regenerating both halves because it cannot tell which one to blame.
3The one script-driven entry publishes the most useful sentence here
Driving lip movement from the spoken content makes a rewrite the only lever, which is a real limitation for drama. The same vendor then says performance improves with the avatar framed closer to camera, which names the weak case.
Naming a weak case is rarer than naming a capability and worth more. Wide framing is the staple of dialogue coverage, so that one sentence constrains a shot list in a way no capability claim does.
4The entries on this route, one page each
Each of these links to that entry on the field this route groups by. The wording behind every cell, and the date it was read, sits on the page it links to.
- Hedra — the supplied audio.
- MiniMax — whoever is on screen.
- sync-3 — the audio it is given.
- Synthesia — the spoken content, framed close.
- Wan2.2-S2V — the audio input.
- Audio sourceAvatar videos are driven by audio, and the character in the image will lip-sync and move to the audio providedhanded to the model rather than made by it
- Audio sourceNative speech with lip-sync tied to the speaker who is on screenspeaker-aware
- Lip-syncLip sync and facial expressions are generated from the spoken contentdriven by the script
- A framing condition attached to lip-syncLip sync performs best when the avatar is framed closer in the scene rather than positioned far from the camera
- Audio sourceThe accepted pairs are video with audio, video with text, image with audio and image with textsupplied or read from text
- A second control alongside the audioA pose video argument lets the result follow a pose sequence while staying synchronised to the audio
5Sources
Membership of this route is decided by the wording in the field it groups by, read from the documentation each vendor publishes. The column itself is on the field note, all five columns are on the speech table, and what counts as documented is on how read. Other routes: Named, not explained, The subject skipped.