Background on why speech is the least documented field
These pages carry the reasoning the notes assume. Speech is documented less than any other part of generative video, and the reasons are technical rather than evasive. As of 2026-09-12.
Lip-sync is the clearest case. Aligning a mouth with a waveform can be done by generating both together, by animating to supplied audio, or by fixing the result afterwards, and the three approaches fail in visibly different ways. A vendor that documents only the feature name has told a production nothing about which failure to expect.
Timing gets its own explainer because it is where editors lose days. A line delivered two seconds longer than the shot forces a choice between recutting the picture and regenerating the audio, and neither is cheap at series length.
The rest divide into three groups. Pages about reading the register itself, including what a blank cell does and does not tell a reader and what one sentence from a vendor would settle. Pages about production decisions the columns force, from coverage and filing discipline to which revision costs a performance. And pages about the sums nobody does, between a shot length, a speaking rate and a subtitle.
1The explainers
- Writing the lines — Generated speech follows the text more literally than an actor does.
- Dub or regenerate — A second language can be reached by replacing the audio or by generating again.
- How lip-sync works — Generated speech can be produced with the picture, laid over it afterwards, or driven from a supplied track.
- Casting a voice — Choosing a generated voice is a casting decision with a constraint attached: whatever you pick has to be reproducible months later, by somebody else.
- Three routes to a voice — A preset library, a cloned voice, or speech generated with the picture.
- Reading a blank cell — An empty cell says a vendor has not published something.
- Counts and lists — A total says a capability is broad.
- Two speakers — The commonest shot in drama is two people with one of them talking, and almost nothing says which face will move.
- Keeping the clip — Where a voice is rebuilt from a reference clip on every call, the archive is what holds a cast together, and the output never records the clip.
- Scheduling a track — Where a recording drives the generation, casting and a session sit in front of the first frame, which reshapes the whole schedule.
- Open weights — On a hosted product an absent sentence is a refusal to commit.
- Two kinds of fix — Where sound and picture come from one call, redoing a shot for a visual reason returns a new performance and an approved line disappears without warning.
- Framing and the mouth — One published condition says lip-sync performs better with a figure framed closer to camera, which turns a mouth problem into a coverage decision.
- Auditioning first — One entry documents a preview.
- Singing and read copy — One vendor names singing and advertisements as supported audio types.
- First tests — The questions that decide a series are the ones documentation answers least, and most of them can be settled with a handful of short generations.
- Subtitles and length — Spoken length and subtitle reading speed are two separate ceilings, and a muted viewer gets the subtitle as the whole performance.
- Asking a vendor — Where documentation is silent, it is worth knowing exactly what a vendor would have to write for the question to become answerable at all.
- Asking for silence — Where audio is on by default, a scene meant to play dry will not, and the fix is a regeneration rather than a mute: the bed is in the file.
- Sizing a scene — A shot with a published ceiling holds a predictable number of words, and the sum takes less time than regenerating a line that overruns.
- A cast list — Only two entries attach a voice to a character object, and even there nothing says how to rebuild one, so the cast list stays a production document.
No page here adds a row to the table; they explain what the rows are worth.
Other notes: Speech controls, Models, Fields, Routes, Side by side, Questions, Terms, Data. What counts as a documented control is set out on the reading page.