Sentioscope

Speech and voice controls, as each vendor documents them

The words that decide how dialogue gets made

Each term here is used loosely in product copy and precisely in these notes, and the gap between those two uses is usually the thing a production ends up planning around. As of 2026-09-12.

Three speech terms, and what each one commits a vendor toNative audio commits to sound from the same model in the same pass. Reference audio commits to following a supplied sample. Line timing commits to a duration a shot can be cut to. Product copy uses all three loosely, and only the first is usually documented precisely.Native audioSound with the pictureSame model, same passUsually documented clearlyReference audioA supplied voice sampleParameters vary widelyDocumented unevenlyLine timingA duration the edit can useRarely named at allDiscovered at the renderWhat a speech feature actually promisesHow well each is documented
Fig. 1 The commitments get weaker from left to right, and the documentation gets thinner in the same direction.

Native audio means sound generated together with the picture, from the same model, in the same pass. A tool that produces silent video and offers a separate voice service is not doing that, however well the two work together in a demo.

Reference audio means supplying a sample so that generated speech follows a particular voice. The term covers arrangements ranging from a few seconds of sample to an enrolled voice profile, and the distinction matters for consent as much as for quality.

The rest of the vocabulary divides into three groups. Words for where the sound comes from, words for how a voice is held to a character, and words for how the result is measured. Each page defines one and logs what the vendors document about it.

1The terms

  • Native audio — Four vendors make speech in the same pass as the shot and publish a clip length with it.
  • Reference audio — Four vendors accept a recording as the source of a voice, with caps on how many and how long.
  • Dialogue timing — Speech runs at a measurable rate and shots have fixed ceilings.
  • Lip-sync — Lip-sync covers three different mechanisms with three different failures, and most vendors name the word without saying which of the three applies.
  • Voice id — A voice id is a short stable reference to a voice, and it is the cheapest continuity mechanism in this register because it cannot drift.
  • Voice cloning — Cloning turns a recording into a reusable voice, and the published thresholds run from one file to twenty minutes or more of studio time.
  • Text to speech — Where speech is synthesised in a separate call, the audio becomes an artefact that can be fetched, approved and reused before any frame exists.
  • Audio-driven video — Three entries treat a supplied waveform as the driving input, which puts casting and a session in front of the first frame and names the driver for free.
  • Speaker label — A speaker label routes a line to a face within one generation, and carries no identity into the next, which is the distinction most often read past.
  • Voice catalogue — A catalogue is a published set of voices to pick from, and the gap between a browsable one and an asserted one decides whether casting is possible.
  • Language code — A language name can mean a market, a script or a dialect.
  • Guide track — Generated speech arrives aligned by construction, which is the hard part, and is usually not a finished mix, which a post budget has to assume.
  • Dubbing pass — A dubbing pass records a translated performance and lays it over finished footage, which is the stage a lip-sync repair exists to clean up.
  • Timbre — Timbre is the colour of a voice, and it is the property that transfers reliably from a very short reference clip.
  • Digital replica — One entry in this register treats cloning an identifiable voice as a likeness question and names United States digital replica statutes beside it.
  • Clip length — A generated shot has a published ceiling, and that ceiling decides how a line has to be written long before anybody spends a call on it.
  • On-screen speaker — With two people in frame, something has to decide whose mouth moves.

Each definition is written to be quoted in a sentence and linked from wherever a note uses the word.

Other notes: Speech controls, Models, Fields, Routes, Side by side, Questions, Learn, Data. What counts as a documented control is set out on the reading page.