Entries that synthesise speech before any picture
Four entries put speech synthesis in front of the picture. A script goes in, audio comes out, and the frames are made against it. Because the audio exists on its own it can be played to whoever signs it off. As of 2026-09-22.
| Model | Where the voice comes from | And what is published about language |
|---|---|---|
| D-ID | A voice id, or a recording by url | A language field, no list |
| HeyGen | A recording to clone from | A count, attached to translation |
| PixVerse | A sample, or a built-in voice | Multiple, none of them named |
| Synthesia | A catalogue voice, or a cloned one | A catalogue with codes |
Inclusion rule. Entries whose documentation describes speech being synthesised from text as a stage, whether at its own endpoint or on the same request. Entries where the sound arrives with the picture are on a separate route page. Order. Alphabetical by model name.
1Splitting the call means a wrong word and a wrong picture cost differently
Where speech is its own step, a misread line costs one speech call, a wrong picture costs the render, and the approved audio survives either. That separation is what productions ask for and rarely get from a single-call generator.
It also gives sign-off something to happen to. An audio file can be circulated, and a producer who has heard the line before frames exist is not going to reject a finished shot for a reason that had nothing to do with the picture.
2These four are the ones with real catalogues
Every entry on this page publishes something a production can choose from, which is not true of any other group here. One names five outside providers, one publishes voices with language codes, one offers built-in and custom voices behind a single id, and one enrols clones at two grades.
That is what a speech stage implies: if speech is a product, the voices have to be selectable, and the vendor has to say something about them. The entries that make sound alongside a picture almost never do.
3A synthesised read is a different craft problem
None of these gives a director the levers a performer gives. A line can be rewritten, a voice swapped and a speed adjusted where documented, and asking for the same words delivered colder is not on any of these pages.
For presentations and explainers that is irrelevant. For drama it is the whole job, and it is the reason this route is common in corporate video and rare in the short-drama pipelines this register is read for.
4The entries on this route, one page each
Each of these links to that entry on the field this route groups by. The wording behind every cell, and the date it was read, sits on the page it links to.
- D-ID — a voice id, or a recording by url.
- HeyGen — a recording to clone from.
- PixVerse — a sample, or a built-in voice.
- Synthesia — a catalogue voice, or a cloned one.
- Audio sourceA script is either text of up to forty thousand characters or an audio url, with audio limited to five minutes for clips and ten for talksa script read aloud or a file
- Voice sourceFive speech providers are named for the voice: Microsoft, ElevenLabs, Amazon, Google and Azure OpenAIoutside catalogues named on the page
- Audio sourceA text to speech endpoint turns a script into speech audio as a step of its ownan endpoint apart from the picture
- Voice sourceA voice clone is instant from one recording or professional from twenty minutes or more, and is then passed as a voice idtwo enrolment routes with a threshold
- Voice sourceText to speech supports both built-in voices and custom voices created from user-provided sample audiocatalogue or supplied sample
- LanguagesEach voice is listed with its formal and native language name, a language code, a gender, a name and a voice idpublished as a catalogue with codes
- Lip-syncLip sync and facial expressions are generated from the spoken contentdriven by the script
5Sources
Membership of this route is decided by the wording in the field it groups by, read from the documentation each vendor publishes. The column itself is on the field note, all five columns are on the speech table, and what counts as documented is on how read. Other routes: Nothing published, A voice that is stored.