Sentioscope

Speech and voice controls, as each vendor documents them

HeyGen: speech is an endpoint, not a by-product

HeyGen's quick start puts speech in an endpoint of its own: a script goes in, speech audio comes out, and the picture is rendered against that. Translation is documented separately, with cloning and lip-sync named together. As of 2026-09-22.

A speech endpoint turns audio into a thing with a nameA script goes to a text to speech endpoint, speech audio comes back, and the picture is rendered against it. Because the audio exists on its own it can be fetched, played to whoever approves it, and reused.ScriptText, edited freelySent to speechCheapest thing to changeVoiceA clone, reused by idNamed on the callEnrolled once, not pershotSpeech audioA file of its ownFetched and approvedSurvives a re-renderOne rendered avatar videoWhat can be inspected before any picture exists
Fig. 1 Splitting the call in two means a wrong word and a wrong picture no longer cost the same to fix.
HeyGen on audio source, statement by statement. Read from the vendor's API quick start on 2026-09-22.
What the documentation settlesWhat it leaves to a take
A text to speech endpoint turns a script into speech audioHow the synthesised read compares with one from a booked session
Video translation is documented for thirty or more languages, with cloning and lip-syncWhich of those languages the cloned voice carries convincingly
A clone is instant from one recording, or professional from twenty minutes or moreWhat the twenty-minute grade buys that the instant one does not

Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.

1An endpoint is a place to stop and listen

Splitting speech out into its own call turns the audio into an artefact with a name. It can be fetched, played to whoever has to approve it, stored beside the script and reused, and none of that requires a frame to have been rendered. For a team whose bottleneck is sign-off rather than throughput, that is the useful shape.

It also means the cost of a change depends on which half changed. A wrong word costs one speech call. A wrong delivery costs the same call again. A wrong picture costs the render and leaves the approved audio alone, which is the separation productions ask for and rarely get.

2The interesting half is that translation carries the count

The language figure on this platform is attached to video translation rather than to the speech endpoint. Read carefully, that is a statement about a workflow: the languages are reachable by translating a finished video, which is a different product from choosing a language before anything is made.

For a series that distinction decides where the localisation stage sits. Translating delivered episodes keeps one master and one performance; generating each language separately would keep none. The documentation describes the former, and the count belongs to it.

3The others that synthesise before they render

Three more entries put a speech stage in front of the picture. They differ on whether that stage is a separate endpoint, a field on the same request, or an implicit part of rendering.

  • D-ID — from a script, or a supplied url.
  • PixVerse — supplied, or read from text.
  • Synthesia — from the script, or uploaded.
  • Audio source
    A text to speech endpoint turns a script into speech audio as a step of its ownan endpoint apart from the pictureHeyGen, API quick start / recorded 2026-09-22
  • Languages
    Video translation is documented for thirty or more languages, with voice cloning and lip-synccounted on the translation endpointHeyGen, API quick start / recorded 2026-09-22
  • Voice source
    A voice clone is instant from one recording or professional from twenty minutes or more, and is then passed as a voice idtwo enrolment routes with a thresholdHeyGen, API quick start / recorded 2026-09-22

4Sources

Read from the API quick start at developers.heygen.com on 2026-09-22. The same column across every entry is on audio source; everything this vendor publishes about speech is on HeyGen. What counts as documented is on how read.