Sentioscope

Speech and voice controls, as each vendor documents them

Speech questions with answers in the documentation

A question is answered here when documentation settles it. How a voice sounds, whether a performance is convincing, and whether a listener would notice a seam are all outside what a page can establish. As of 2026-09-12.

What documentation settles about speech, and what only listening canWhether audio is produced, whether a reference is accepted and which languages are named are published facts. Whether a voice is convincing, whether an accent is right for a market and whether a seam is audible are judgements that need ears and a specific script.Documentation settlesOnly listening settlesIs audio producedYes, by a published sentenceCan a voice be suppliedYes, with the parametersHow closely it is followedWhich languagesThe names on the pageWhether the accent suits the marketLine timingWhether control existsWhether the result sounds rushedTwo halves of every speech question
Fig. 1 The right-hand column is where most of the value of a speech feature sits, and no page can establish it.

The exclusion is unusually painful for this subject, because speech is judged by ear and the ear is exactly what a register cannot use. A note can say that a model accepts a reference voice; only listening reveals how closely it follows one.

What documentation does settle turns out to shape a production anyway. Whether audio arrives with the picture decides whether a dubbing stage exists. Whether a language is named decides whether a market is reachable. Whether timing is controllable decides how the edit is planned.

Answers link to the notes, and the notes link to the vendor pages, so a claim can be traced to the sentence a vendor actually published.

1The questions

  • Keeping a voice — Four models keep a voice as a stored object you approve once.
  • Lip-sync — One vendor describes a mechanism, one names the feature without describing it, one removes the step, and several take the audio in so sync is the task.
  • Voice sources — A catalogue, a clone from a recording, a finished track, or a voice described in words.
  • Sound by default — Two vendors state that sound comes back unless it is switched off, and one of them publishes what switching it off does to the price.
  • Which language — One vendor publishes a catalogue with language codes, one names fifteen dialogue languages, six publish a count or a claim, and nine publish nothing.
  • Two in frame — Two characters in frame with a line each is the commonest shot in drama, and exactly one vendor publishes any guidance for it.
  • Who docs are for — Developer references describe parameters.
  • How long a shot can be — Eight entries publish a length for a generation or its audio, from a four-second shot to a ten-minute maximum.
  • Fifteen seconds of voice — Two entries cap reference audio at fifteen seconds in total, and the published windows elsewhere run from one recording to twenty minutes of studio time.
  • Answers software can check — Most of this register is prose.
  • What silence costs — One entry states plainly that turning audio off does not change the rate.
  • Where a line goes — Four entries tell a writer where dialogue belongs: a labelled block, a clause after one keyword, a quoted line, or a reference voice instead of words.
  • Auditioning a voice — One entry documents a preview on a voice that has reached a ready state.
  • Outside suppliers named — Two entries name outside companies for speech: one lists five providers behind its voice field, and one names an audio house as a planned integration.
  • The column always filled — Audio source is filled by fifteen of the seventeen entries.
  • Speech without the bed — One entry documents a speech-only mode.
  • Regenerating a shot — Where audio comes out with the picture, redoing a shot produces a new performance.
  • Consent in the docs — One entry treats cloning an identifiable voice as a likeness question and names digital replica statutes.
  • A second control — Four entries publish a control most of the register lacks: a pose track, a speech-only mode, a reverse audio route, or a voice asked for in words.
  • What kind of page — The shape of an answer follows the kind of page it sits on: an API reference, a model card, a prompting guide or a compliance checklist.
  • Formats and file sizes — Four entries publish what they accept as a file: WAV or MP3, a ten-megabyte ceiling, or a hundred megabytes on both sides of one call.
  • What sets the length — Length comes from three places: a menu of fixed options, a range a caller chooses, or the recording a production supplies.
  • Charging for audio — No entry here bills audio as a separate line.

Where documentation is silent, the answer says which sentence would have to appear for the question to become answerable.

Other notes: Speech controls, Models, Fields, Routes, Side by side, Terms, Learn, Data. What counts as a documented control is set out on the reading page.