Sentioscope

Speech and voice controls, as each vendor documents them

Background on why speech is the least documented field

These pages carry the reasoning the notes assume. Speech is documented less than any other part of generative video, and the reasons are technical rather than evasive. As of 2026-09-12.

Where a dialogue pipeline breaks, and who notices firstA mouth that drifts is noticed by a viewer. A line that overruns is noticed by an editor. A voice that changes between episodes is noticed by an audience halfway through a season. The three failures happen at different stages and are documented in inverse order to how much they cost.Noticed soonest at the topLip movement driftsVisible in a single clip. The failure a demo would expose, andthe one vendors address.A line overruns the shotFound at the edit. Costs a recut or a regeneration, per line.A voice changes acrossepisodesFound by the audience. Costs the illusion the whole season restson.Documented least at the bottom
Fig. 1 The cheapest failure to document is the one a demo would reveal anyway, which is part of why the others stay unaddressed.

Lip-sync is the clearest case. Aligning a mouth with a waveform can be done by generating both together, by animating to supplied audio, or by fixing the result afterwards, and the three approaches fail in visibly different ways. A vendor that documents only the feature name has told a production nothing about which failure to expect.

Timing gets its own explainer because it is where editors lose days. A line delivered two seconds longer than the shot forces a choice between recutting the picture and regenerating the audio, and neither is cheap at series length.

The rest divide into three groups. Pages about reading the register itself, including what a blank cell does and does not tell a reader and what one sentence from a vendor would settle. Pages about production decisions the columns force, from coverage and filing discipline to which revision costs a performance. And pages about the sums nobody does, between a shot length, a speaking rate and a subtitle.

1The explainers

  • Writing the lines — Generated speech follows the text more literally than an actor does.
  • Dub or regenerate — A second language can be reached by replacing the audio or by generating again.
  • How lip-sync works — Generated speech can be produced with the picture, laid over it afterwards, or driven from a supplied track.
  • Casting a voice — Choosing a generated voice is a casting decision with a constraint attached: whatever you pick has to be reproducible months later, by somebody else.
  • Three routes to a voice — A preset library, a cloned voice, or speech generated with the picture.
  • Reading a blank cell — An empty cell says a vendor has not published something.
  • Counts and lists — A total says a capability is broad.
  • Two speakers — The commonest shot in drama is two people with one of them talking, and almost nothing says which face will move.
  • Keeping the clip — Where a voice is rebuilt from a reference clip on every call, the archive is what holds a cast together, and the output never records the clip.
  • Scheduling a track — Where a recording drives the generation, casting and a session sit in front of the first frame, which reshapes the whole schedule.
  • Open weights — On a hosted product an absent sentence is a refusal to commit.
  • Two kinds of fix — Where sound and picture come from one call, redoing a shot for a visual reason returns a new performance and an approved line disappears without warning.
  • Framing and the mouth — One published condition says lip-sync performs better with a figure framed closer to camera, which turns a mouth problem into a coverage decision.
  • Auditioning first — One entry documents a preview.
  • Singing and read copy — One vendor names singing and advertisements as supported audio types.
  • First tests — The questions that decide a series are the ones documentation answers least, and most of them can be settled with a handful of short generations.
  • Subtitles and length — Spoken length and subtitle reading speed are two separate ceilings, and a muted viewer gets the subtitle as the whole performance.
  • Asking a vendor — Where documentation is silent, it is worth knowing exactly what a vendor would have to write for the question to become answerable at all.
  • Asking for silence — Where audio is on by default, a scene meant to play dry will not, and the fix is a regeneration rather than a mute: the bed is in the file.
  • Sizing a scene — A shot with a published ceiling holds a predictable number of words, and the sum takes less time than regenerating a line that overruns.
  • A cast list — Only two entries attach a voice to a character object, and even there nothing says how to rebuild one, so the cast list stays a production document.

No page here adds a row to the table; they explain what the rows are worth.

Other notes: Speech controls, Models, Fields, Routes, Side by side, Questions, Terms, Data. What counts as a documented control is set out on the reading page.