Text to speech: making the audio a stage of its own
Text to speech here means a call that takes a script and returns speech audio, before any picture is made. Four entries put that stage in front of the render, and it changes what a revision costs. As of 2026-09-12.
| Model | Where the stage sits | What it produces |
|---|---|---|
| D-ID | Inside the talk request | Speech read from a script of up to 40,000 characters |
| HeyGen | Its own endpoint | Speech audio, fetched and approved separately |
| PixVerse | Inside the speech and lip sync endpoint | Audio from a built-in or custom voice |
| Synthesia | Implicit in the render | Lip sync and expression from the same script |
Inclusion rule. Entries whose documentation describes speech being synthesised from text as part of producing video. Entries where sound arrives with the picture from one generation are a different arrangement and do not earn a row. Order. Alphabetical by model name.
1An artefact is something sign-off can happen to
An audio file has a name, can be circulated and can be played to whoever must approve it. None of that requires a frame, which suits a team whose bottleneck is approval rather than throughput.
It also splits the cost of a change. A wrong word costs one speech call, a wrong delivery costs the same call again, and a wrong picture costs the render and leaves approved audio alone. That separation is what productions ask for and rarely get.
2A separate stage forces a real catalogue
If speech is a product, the voices have to be selectable, so a vendor has to publish something about them. Every entry with a speech stage publishes a catalogue, a clone route or a named supplier. Entries that make sound alongside a picture almost never do.
That correlation is one of the clearer patterns in this register, and it is architectural rather than editorial. A voice nobody selects needs no documentation.
3What a synthesised read cannot be asked for
None of these gives a director the levers a performer gives. A line can be rewritten, a voice swapped and sometimes a speed set. Asking for the same words colder, or faster, or through clenched teeth is not on any of these pages.
For explainers and presentations that is irrelevant, which is why this arrangement is common in corporate video. For drama it is the whole job, and the gap is the reason the route is rare in short-drama pipelines.
4Sources read for this entry
This page defines the term and logs what each vendor documents about it, each figure read on 2026-09-12. Related: Audio-driven video, Speaker label.