HeyGen: two grades of clone, with a threshold in minutes
A clone is instant from a single recording, or professional from twenty minutes or more of audio, and afterwards it travels as a voice id. Publishing the threshold in minutes is unusual: most vendors describe cloning without saying what it costs in studio time. As of 2026-09-22.
| What the documentation settles | What it leaves to a take |
|---|---|
| A clone is instant from one recording, or professional from twenty minutes or more | What the professional grade actually sounds better at |
| The clone is then passed as a voice id in the request | Whether both grades behave identically once they have an id |
| A text to speech endpoint turns a script into speech audio | Whether the clone is available to that endpoint as well as to video |
Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.
1A threshold in minutes is a production decision, not a parameter
Twenty minutes of usable audio from one performer is a booked session, a quiet room and a release form. Publishing that figure turns voice cloning from a feature into a line in a schedule, which is the most useful thing a page can do with it.
The instant grade is the counterpart and the trap. It works from whatever recording is at hand, which makes it easy to try and easy to commit to by accident, and a season built on a clone taken from a phone memo is expensive to re-voice later.
2Two grades are two different consent conversations
A performer who sits for twenty minutes knows what is being built. A recording reused from elsewhere may not carry that understanding, and the difference matters more as the clone travels further: this platform documents cloning alongside translation into thirty or more languages.
Nothing on the page addresses permission, which is normal for an API quick start and worth noting anyway. The register records the mechanism; who may lawfully be cloned is decided somewhere else entirely.
3The other entry where a voice is built once and reused
One more entry treats a voice as something constructed before any shot exists and afterwards addressed by name. It offers a route this one does not: a voice asked for in words.
- Runway — a 10-second to 5-minute sample, or a sentence.
- Voice sourceA voice clone is instant from one recording or professional from twenty minutes or more, and is then passed as a voice idtwo enrolment routes with a threshold
- Audio sourceA text to speech endpoint turns a script into speech audio as a step of its ownan endpoint apart from the picture
4Sources
Read from the API quick start at developers.heygen.com on 2026-09-22. The same column across every entry is on voice source; everything this vendor publishes about speech is on HeyGen. What counts as documented is on how read.