HeyGen: speech is an endpoint, not a by-product
HeyGen's quick start puts speech in an endpoint of its own: a script goes in, speech audio comes out, and the picture is rendered against that. Translation is documented separately, with cloning and lip-sync named together. As of 2026-09-22.
| What the documentation settles | What it leaves to a take |
|---|---|
| A text to speech endpoint turns a script into speech audio | How the synthesised read compares with one from a booked session |
| Video translation is documented for thirty or more languages, with cloning and lip-sync | Which of those languages the cloned voice carries convincingly |
| A clone is instant from one recording, or professional from twenty minutes or more | What the twenty-minute grade buys that the instant one does not |
Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.
1An endpoint is a place to stop and listen
Splitting speech out into its own call turns the audio into an artefact with a name. It can be fetched, played to whoever has to approve it, stored beside the script and reused, and none of that requires a frame to have been rendered. For a team whose bottleneck is sign-off rather than throughput, that is the useful shape.
It also means the cost of a change depends on which half changed. A wrong word costs one speech call. A wrong delivery costs the same call again. A wrong picture costs the render and leaves the approved audio alone, which is the separation productions ask for and rarely get.
2The interesting half is that translation carries the count
The language figure on this platform is attached to video translation rather than to the speech endpoint. Read carefully, that is a statement about a workflow: the languages are reachable by translating a finished video, which is a different product from choosing a language before anything is made.
For a series that distinction decides where the localisation stage sits. Translating delivered episodes keeps one master and one performance; generating each language separately would keep none. The documentation describes the former, and the count belongs to it.
3The others that synthesise before they render
Three more entries put a speech stage in front of the picture. They differ on whether that stage is a separate endpoint, a field on the same request, or an implicit part of rendering.
- D-ID — from a script, or a supplied url.
- PixVerse — supplied, or read from text.
- Synthesia — from the script, or uploaded.
- Audio sourceA text to speech endpoint turns a script into speech audio as a step of its ownan endpoint apart from the picture
- LanguagesVideo translation is documented for thirty or more languages, with voice cloning and lip-synccounted on the translation endpoint
- Voice sourceA voice clone is instant from one recording or professional from twenty minutes or more, and is then passed as a voice idtwo enrolment routes with a threshold
4Sources
Read from the API quick start at developers.heygen.com on 2026-09-22. The same column across every entry is on audio source; everything this vendor publishes about speech is on HeyGen. What counts as documented is on how read.