Audio-driven video: the recording as the instruction
Audio-driven means the sound is the input and the picture is fitted to it. Three entries work this way. Nothing is generated until a recording exists, and all three name what the mouth is following. As of 2026-09-12.
| Model | What it accepts | What else can ride alongside |
|---|---|---|
| Hedra | A track, or a voice id generated inline | The recording normally sets the length |
| sync-3 | Audio or text, with footage or a still | Nothing else is published |
| Wan2.2-S2V | An audio input with a reference image | An optional pose video sequence |
Inclusion rule. Entries whose documentation describes audio as an input that drives generation. Entries that accept a short reference clip to steer a voice they then generate are a different arrangement and do not earn a row. Order. Alphabetical by model name.
1The driver comes free, and so does the approval
When the waveform is an input, saying the mouth follows it is a description of the interface rather than a promise about a model's internals. All three publish it, while only two entries elsewhere in this register manage a driver statement at all.
A director also signs off the performance before any generation is billed, which is the strongest argument for this arrangement. What comes back afterwards is a picture problem rather than a performance one.
2The schedule runs backwards, and a line change is a re-record
Casting, a session and a delivered file all sit in front of the first frame, which lengthens the front of a schedule and shortens the part involving rendering. The compensation is that a visual fix leaves the approved read untouched.
Changing a line, though, costs a room, a performer and a diary rather than a call. Whether that is a good trade depends on whether dialogue or art direction moves more often on a given show.
3Two of the three are stages rather than generators
Products here operate on material: a still, a reference image or existing footage. They sit late in a pipeline and assume something else produced the picture, which is a scheduling fact rather than a limitation.
It also explains why voice inventories are thin on this route. A team reaching this stage usually has its audio already, from a session or from whichever speech service it standardised on, so a catalogue would go unused.
4Sources read for this entry
This page defines the term and logs what each vendor documents about it, each figure read on 2026-09-12. Related: Speaker label, Voice catalogue.