Sentioscope

Speech and voice controls, as each vendor documents them

Text to speech: making the audio a stage of its own

Text to speech here means a call that takes a script and returns speech audio, before any picture is made. Four entries put that stage in front of the render, and it changes what a revision costs. As of 2026-09-12.

Splitting the call so revisions cost differentlyWhere speech is its own stage, a wrong word costs one speech call, a wrong delivery the same call again, and a wrong picture the render alone. The approved audio survives any of the three.ScriptText, edited freelyInto a speech callCheapest to changeSpeech audioA file with a nameFetched and approvedSurvives a re-renderPictureRendered against itRegenerated aloneLeaves the audio aloneOne delivered videoWhat nothing here provides: direction on the reading
Fig. 1 If speech is a product, the voices have to be selectable, which is why every entry on this route publishes a catalogue of some kind.
Text to speech, as published. Recorded 2026-09-12.
ModelWhere the stage sitsWhat it produces
D-IDInside the talk requestSpeech read from a script of up to 40,000 characters
HeyGenIts own endpointSpeech audio, fetched and approved separately
PixVerseInside the speech and lip sync endpointAudio from a built-in or custom voice
SynthesiaImplicit in the renderLip sync and expression from the same script

Inclusion rule. Entries whose documentation describes speech being synthesised from text as part of producing video. Entries where sound arrives with the picture from one generation are a different arrangement and do not earn a row. Order. Alphabetical by model name.

1An artefact is something sign-off can happen to

An audio file has a name, can be circulated and can be played to whoever must approve it. None of that requires a frame, which suits a team whose bottleneck is approval rather than throughput.

It also splits the cost of a change. A wrong word costs one speech call, a wrong delivery costs the same call again, and a wrong picture costs the render and leaves approved audio alone. That separation is what productions ask for and rarely get.

2A separate stage forces a real catalogue

If speech is a product, the voices have to be selectable, so a vendor has to publish something about them. Every entry with a speech stage publishes a catalogue, a clone route or a named supplier. Entries that make sound alongside a picture almost never do.

That correlation is one of the clearer patterns in this register, and it is architectural rather than editorial. A voice nobody selects needs no documentation.

3What a synthesised read cannot be asked for

None of these gives a director the levers a performer gives. A line can be rewritten, a voice swapped and sometimes a speed set. Asking for the same words colder, or faster, or through clenched teeth is not on any of these pages.

For explainers and presentations that is irrelevant, which is why this arrangement is common in corporate video. For drama it is the whole job, and the gap is the reason the route is rare in short-drama pipelines.

4Sources read for this entry

This page defines the term and logs what each vendor documents about it, each figure read on 2026-09-12. Related: Audio-driven video, Speaker label.