Sentioscope

Speech and voice controls, as each vendor documents them

Entries that synthesise speech before any picture

Four entries put speech synthesis in front of the picture. A script goes in, audio comes out, and the frames are made against it. Because the audio exists on its own it can be played to whoever signs it off. As of 2026-09-22.

Why a speech stage forces a real catalogueIf speech is a product of its own, the voices have to be selectable, so the vendor has to publish something about them. Every entry on this route does, and the entries that make sound alongside a picture almost never do.ScriptText, cheap to editSent to speechRewrites cost almostnothingVoiceChosen from a catalogueNamed on the callSelectable, because it isa productSpeech audioA file of its ownApproved, then renderedSurvives a re-renderOne rendered videoWhat exists before any frame does
Fig. 1 A misread line costs one speech call, a wrong picture costs the render, and the approved audio survives either.
Script-to-speech entries, beside the voice route and language answer each publishes. Recorded 2026-09-22.
ModelWhere the voice comes fromAnd what is published about language
D-IDA voice id, or a recording by urlA language field, no list
HeyGenA recording to clone fromA count, attached to translation
PixVerseA sample, or a built-in voiceMultiple, none of them named
SynthesiaA catalogue voice, or a cloned oneA catalogue with codes

Inclusion rule. Entries whose documentation describes speech being synthesised from text as a stage, whether at its own endpoint or on the same request. Entries where the sound arrives with the picture are on a separate route page. Order. Alphabetical by model name.

1Splitting the call means a wrong word and a wrong picture cost differently

Where speech is its own step, a misread line costs one speech call, a wrong picture costs the render, and the approved audio survives either. That separation is what productions ask for and rarely get from a single-call generator.

It also gives sign-off something to happen to. An audio file can be circulated, and a producer who has heard the line before frames exist is not going to reject a finished shot for a reason that had nothing to do with the picture.

2These four are the ones with real catalogues

Every entry on this page publishes something a production can choose from, which is not true of any other group here. One names five outside providers, one publishes voices with language codes, one offers built-in and custom voices behind a single id, and one enrols clones at two grades.

That is what a speech stage implies: if speech is a product, the voices have to be selectable, and the vendor has to say something about them. The entries that make sound alongside a picture almost never do.

3A synthesised read is a different craft problem

None of these gives a director the levers a performer gives. A line can be rewritten, a voice swapped and a speed adjusted where documented, and asking for the same words delivered colder is not on any of these pages.

For presentations and explainers that is irrelevant. For drama it is the whole job, and it is the reason this route is common in corporate video and rare in the short-drama pipelines this register is read for.

4The entries on this route, one page each

Each of these links to that entry on the field this route groups by. The wording behind every cell, and the date it was read, sits on the page it links to.

  • D-ID — a voice id, or a recording by url.
  • HeyGen — a recording to clone from.
  • PixVerse — a sample, or a built-in voice.
  • Synthesia — a catalogue voice, or a cloned one.
  • Audio source
    A script is either text of up to forty thousand characters or an audio url, with audio limited to five minutes for clips and ten for talksa script read aloud or a fileD-ID, create a talk reference / recorded 2026-09-22
  • Voice source
    Five speech providers are named for the voice: Microsoft, ElevenLabs, Amazon, Google and Azure OpenAIoutside catalogues named on the pageD-ID, create a talk reference / recorded 2026-09-22
  • Audio source
    A text to speech endpoint turns a script into speech audio as a step of its ownan endpoint apart from the pictureHeyGen, API quick start / recorded 2026-09-22
  • Voice source
    A voice clone is instant from one recording or professional from twenty minutes or more, and is then passed as a voice idtwo enrolment routes with a thresholdHeyGen, API quick start / recorded 2026-09-22
  • Voice source
    Text to speech supports both built-in voices and custom voices created from user-provided sample audiocatalogue or supplied samplePixVerse, speech and lip sync guide / recorded 2026-09-22
  • Languages
    Each voice is listed with its formal and native language name, a language code, a gender, a name and a voice idpublished as a catalogue with codesSynthesia, list of supported voices / recorded 2026-09-22
  • Lip-sync
    Lip sync and facial expressions are generated from the spoken contentdriven by the scriptSynthesia, create an avatar / recorded 2026-09-22

5Sources

Membership of this route is decided by the wording in the field it groups by, read from the documentation each vendor publishes. The column itself is on the field note, all five columns are on the speech table, and what counts as documented is on how read. Other routes: Nothing published, A voice that is stored.