HeyGen: lip-sync named, inside a translation sentence
Lip-sync appears in one place: the sentence describing video translation for thirty or more languages, with voice cloning beside it. It is named as part of a workflow rather than explained as a mechanism. As of 2026-09-22.
| What the documentation settles | What it leaves to a take |
|---|---|
| Video translation is documented for thirty or more languages, with cloning and lip-sync | What drives the mouth during that translation |
| A text to speech endpoint turns a script into speech audio | Whether lip-sync applies to speech from that endpoint too |
| A clone is instant from one recording, or professional from twenty minutes or more | Whether a cloned voice aligns better than a catalogue one |
Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.
1Where a claim appears decides what it covers
Because lip-sync is mentioned only in the translation sentence, the safe reading is that it is documented for translation. Whether the same behaviour applies when a video is generated from a script in its original language is not stated, and assuming it does would be extending the sentence past what it says.
That is a narrow reading and a deliberate one. This column has three grades of answer, and a claim attached to one workflow is worth less than a driver named for the product as a whole.
2Translating a mouth is a harder problem than generating one
Translation means new words of a different length arriving over footage that already exists. The mouth has to be changed rather than created, and the surrounding face has to stay put. Naming lip-sync in exactly that sentence is therefore a claim about the harder case.
It is also the case where failure is most visible, because a viewer has the original performance's rhythm to compare against. Nothing published says how the trade is made, so the cell records the claim and not a mechanism.
3The others that name the feature and stop
Three more entries put lip-sync on the page without saying what moves the mouth. One of them names it only in a warning about where it stops holding.
- Kling AI — named, with nothing driving it.
- PixVerse — named as the endpoint's purpose.
- Sora 2 — named only where it fails.
- LanguagesVideo translation is documented for thirty or more languages, with voice cloning and lip-synccounted on the translation endpoint
- Voice sourceA voice clone is instant from one recording or professional from twenty minutes or more, and is then passed as a voice idtwo enrolment routes with a threshold
- Audio sourceA text to speech endpoint turns a script into speech audio as a step of its ownan endpoint apart from the picture
4Sources
Read from the API quick start at developers.heygen.com on 2026-09-22. The same column across every entry is on lip-sync; everything this vendor publishes about speech is on HeyGen. What counts as documented is on how read.