Scheduling around audio that has to exist first
An audio-driven pipeline reverses the usual order. Nothing can be generated until a performance exists, so the risky, slow work happens before any rendering is paid for and a director approves the read first. As of 2026-09-12.
| Stage | Sound generated | Sound supplied |
|---|---|---|
| Approval of the read | After a render, if at all | Before any frame exists |
| Casting | Effectively skipped | A real stage with a diary |
| Cost of a line change | A generation | A room and a performer |
| Cost of a picture change | A new performance too | One generation |
Inclusion rule. Stages of a dialogue production, compared between the two documented architectures. Order. Alphabetical by stage.
1The slow part moves to the front, where it is cheaper
Booking a performer and a room is slower than making a call, and it is far cheaper to change. A line rewritten before a session costs a keystroke; the same line rewritten after delivery costs a booking.
So an audio-driven schedule looks worse on paper and behaves better under revision, provided the script is close to locked. Where dialogue is still moving, the arrangement is the wrong way round.
2Pacing becomes a decision taken with a microphone present
Where the recording determines the length, a performer who pauses for effect is buying frames. That is the cheapest possible place to make a pacing decision and the most awkward place to revise one.
A production working to fixed episode lengths therefore has to bring the timing requirement into the session, rather than discovering at the edit that a scene runs eight seconds long.
3Two of these products are stages, not generators
Audio-driven tools mostly operate on material: a still, a reference image or existing footage. Something upstream has to produce the picture, and that dependency belongs on the schedule as its own line.
It also explains why they publish little about voices. A team reaching this stage usually has its audio already, so a catalogue would go unused and goes undocumented.
4Where the published detail is logged
What is described above is how the techniques work in general. Which models document which of them, in whose words, is kept in the speech controls table with a date on every field.
- Entries handed a recording — the three that work this way
- Audio-driven video — the recording as the instruction
A mechanism note rather than a documented field. Nothing here is attributed to a model, and nothing here is a claim about one. The sourced material is on the speech controls table. Related: Open weights, Two kinds of fix.