What HeyGen documents about speech
HeyGen (developers.heygen.com) documents a voice clone that is instant from one recording or professional from twenty minutes or more and is then passed by id, plus video translation into thirty or more languages with lip-sync. As of 2026-09-22.
| Field | What the vendor documents |
|---|---|
| Audio source | A text to speech endpoint, apart from the picture |
| Languages | Thirty or more, counted on the translation endpoint |
| Voice source | A clone, instant from one recording or professional from 20 minutes or more |
| Per-character binding | A cloned voice, reused by voice id |
| Lip-sync | Documented as part of video translation |
Inclusion rule. Fields are the same five for every model here; a field the vendor does not document is recorded as not documented rather than inferred from generated output. Order. Fixed field order, identical on every model page.
1Two grades of the same voice, with the threshold published
An instant clone comes from a single recording; a professional one asks for twenty minutes or more. Publishing both, with the boundary between them stated in minutes, tells a production what it has to book: a phone memo for a background part, a studio hour for a lead.
Once made, the clone is an id passed in the request, which makes the voice a reusable asset rather than a per-shot upload. That is the arrangement a season needs, and this entry states it in the same breath as the enrolment cost.
2Languages are counted where translation happens
The language figure on this platform is attached to translation rather than to generation: an existing video is translated and dubbed into thirty or more languages, with voice cloning and lip-sync. A production reading that as a list of languages it can perform in from scratch would be stretching the sentence.
Read as written it describes a route into another market that begins from a finished episode. That is the dubbing path with the mouth repaired afterwards, and it is a different plan from generating the performance in the target language to begin with.
3Speech as a step, and what that does to a schedule
Text becomes speech audio at its own endpoint, and the avatar is rendered from that. Pulling the two apart is useful: a line can be approved before a frame is rendered, and a re-render for a visual reason need not produce a new performance.
It also means nothing here arrives in one pass, so alignment is a service rather than a by-product. The register puts the lip-sync claim where the vendor makes it, beside translation, and leaves the mechanism unstated because the page read does not give one.
4This entry, one column at a time
Each of these stays inside a single field: what HeyGen puts there, what the wording settles, and what it leaves for a take to answer.
- Audio source — from a speech endpoint, then rendered.
- Voice source — a recording to clone from.
- Per-character binding — on a stored clone.
- Languages — a count, attached to translation.
- Lip-sync — named, alongside translation.
5Read against another entry
Each of these puts HeyGen beside one other entry on a column where the two land at opposite grades of answer.
- HeyGen and LTX Studio — on per-character binding.
- HeyGen and Synthesia — on languages.
6Where this model sits against the rest
| Model | Audio source |
|---|---|
| D-ID | A text script read aloud, or an audio url supplied |
| Hedra | Supplied; audio is a required input |
| HeyGen | A text to speech endpoint, apart from the picture |
| Kling AI | Generated with the picture, VIDEO 3.0 |
| LTX Studio | Generated jointly with the picture, plus audio-to-video |
| Luma Ray | Not documented; the video reference does not mention sound |
| MiniMax | Native speech, generated with the picture |
| PixVerse | Supplied to a speech endpoint, or read from text there |
| Runway | Not documented where voices are defined |
| SceneMixer | Generated with the picture, in the language set for the project |
| Sora 2 | Generated with the picture; video and audio are both listed as output |
| sync-3 | Supplied, or read from text on the call |
| Synthesia | Synthesised from the script, or an uploaded recording |
| Veo | Generated with the video rather than added afterwards |
| Vidu | Q3 generates speech, effects and music natively, on by default |
| Wan 3.0 | On by default; false returns a file with no audio track |
| Wan2.2-S2V | Supplied; the card describes the model as audio-driven |
Inclusion rule. Models whose vendor documents something about generated speech. A model with nothing published on the point is still listed, with the cell marked hollow. Order. Alphabetical by model name.
| Model | Languages |
|---|---|
| D-ID | A language field beside the voice; no list on the reference |
| Hedra | Full multi-language support claimed, none named |
| HeyGen | Thirty or more, counted on the translation endpoint |
| Kling AI | Five, not named |
| LTX Studio | Not documented |
| Luma Ray | Not documented |
| MiniMax | Not documented by the vendor |
| PixVerse | Multiple, with speech, singing and advertisements named as types |
| Runway | Not documented |
| SceneMixer | 15, named, with Cantonese a 16th for dialogue only |
| Sora 2 | Not documented |
| sync-3 | Ninety-five or more, counted rather than named |
| Synthesia | A catalogue, each voice carrying a language name and code |
| Veo | Not documented |
| Vidu | Not documented |
| Wan 3.0 | Not documented |
| Wan2.2-S2V | Not documented |
Inclusion rule. Models whose vendor documents something about generated speech. A model with nothing published on the point is still listed, with the cell marked hollow. Order. Alphabetical by model name.
| Model | Voice source |
|---|---|
| D-ID | A voice id from one of five named speech providers |
| Hedra | An uploaded track, or a voice id from the voices endpoint |
| HeyGen | A clone, instant from one recording or professional from 20 minutes or more |
| Kling AI | Not documented by the vendor |
| LTX Studio | Not documented |
| Luma Ray | Not documented |
| MiniMax | Reference audio, 15 s total across 3 clips |
| PixVerse | Built-in voices, or a custom voice from a supplied sample |
| Runway | An audio sample of 10 seconds to 5 minutes, or a description in words |
| SceneMixer | Preset voice library, or the user's own recordings |
| Sora 2 | Not documented |
| sync-3 | Not documented on the model page |
| Synthesia | The catalogue, or a cloned voice with a language list of its own |
| Veo | Not documented |
| Vidu | Not documented |
| Wan 3.0 | Reference audio, 15 seconds in total, WAV or MP3 up to 15 MB |
| Wan2.2-S2V | Whatever track is handed over; no catalogue |
Inclusion rule. Models whose vendor documents something about generated speech. A model with nothing published on the point is still listed, with the cell marked hollow. Order. Alphabetical by model name.
7Sources
Taken from the API quick start at developers.heygen.com on 2026-09-22. Every entry against the same five fields sits on the speech controls page; the rule behind each field is on how read. The entries either side of this one: Kling AI, LTX Studio.
- Voice sourceA voice clone is instant from one recording or professional from twenty minutes or more, and is then passed as a voice idtwo enrolment routes with a threshold
- LanguagesVideo translation is documented for thirty or more languages, with voice cloning and lip-synccounted on the translation endpoint
- Audio sourceA text to speech endpoint turns a script into speech audio as a step of its ownan endpoint apart from the picture