Entries handed a recording and asked to fit a picture
Three entries are driven by audio a production supplies. Nothing is generated until a recording exists, which lengthens the front of a schedule and hands a director something to approve before any rendering is paid for. As of 2026-09-22.
| Model | What the model is given | And what the mouth is said to follow |
|---|---|---|
| Hedra | Supplied; audio is required | The supplied audio |
| sync-3 | Supplied, or read from text | The audio it is given |
| Wan2.2-S2V | Supplied; the model is audio-driven | The audio input |
Inclusion rule. Entries whose documentation describes audio as an input that drives the generation. Entries that synthesise speech from a script in a separate step are on their own route page, because the recording session is the part that differs. Order. Alphabetical by model name.
1Every entry here names its driver, and that is not a coincidence
When the waveform is an input, saying the mouth follows it costs the vendor nothing: it is a description of the interface rather than a promise about a model's internals. All three of these publish a driver statement, and only two entries elsewhere in the register manage one.
That makes this group the easiest to plan around and the hardest to shortcut. The alignment question is settled on the page; the recording question is entirely a production's own, and no parameter will help with it.
2Approval moves to the front, and so does the cost of a change
A director signs off a performance before a frame is billed, which is the strongest argument for this arrangement. The counterpart is that a line change is a re-record, and a re-record is a room, a performer and a diary rather than a call.
One of these says the audio normally determines the video length, which pushes even the pacing decision into the session. A performer who pauses for effect is buying frames, and that is much cheaper to adjust with a microphone in front of them.
3Two of the three are stages rather than generators
Products on this page operate on material: a still, a reference image, or existing footage. They sit late in a pipeline and assume something else made the picture or the portrait, which is a scheduling fact rather than a limitation.
It also explains why voice inventories are thin here. A team reaching this stage usually already has its audio, from a session or from whichever speech service it standardised on, so a catalogue would go unused.
4The entries on this route, one page each
Each of these links to that entry on the field this route groups by. The wording behind every cell, and the date it was read, sits on the page it links to.
- Hedra — supplied; audio is required.
- sync-3 — supplied, or read from text.
- Wan2.2-S2V — supplied; the model is audio-driven.
- Audio sourceAvatar videos are driven by audio, and the character in the image will lip-sync and move to the audio providedhanded to the model rather than made by it
- Shot length set by the takeThe audio normally determines the video length, and omitting the duration follows the source audio
- Audio sourceThe accepted pairs are video with audio, video with text, image with audio and image with textsupplied or read from text
- Length per call on the free tierA free account runs one generation a month with a fifteen-second ceiling, with paid limits following the plan
- Audio sourceThe card describes audio-driven cinematic video generation from an audio input with a reference image and an optional text prompthanded to the model rather than made by it
- A second control alongside the audioA pose video argument lets the result follow a pose sequence while staying synchronised to the audio
5Sources
Membership of this route is decided by the wording in the field it groups by, read from the documentation each vendor publishes. The column itself is on the field note, all five columns are on the speech table, and what counts as documented is on how read. Other routes: Speech from a script, Nothing published.