Voice cloning: building a voice from a recording
Cloning builds a reusable voice from audio of a real person. Four entries document it, and the amount of audio asked for is the figure that decides whether it is a feature or a booked session. As of 2026-09-12.
| Model | What it asks for | What the clone becomes |
|---|---|---|
| HeyGen | One recording, or twenty minutes and up | A voice id, usable in translation too |
| PixVerse | A supplied sample, length unpublished | A speaker id on the request |
| Runway | Ten seconds to five minutes, up to 10 MB | A stored voice with a preview |
| Synthesia | A recording, with its own language range | A voice alongside the catalogue |
Inclusion rule. Entries whose documentation describes constructing a reusable voice from supplied audio. A reference clip that establishes a voice for one call only is a different mechanism and does not earn a row. Order. Alphabetical by model name.
1A published threshold turns a feature into a schedule line
Twenty minutes of usable audio from one performer means a booked session, a quiet room and a release form. Publishing that figure is the most useful thing a vendor can do with cloning, because it moves the decision from a product page onto a calendar.
The instant grade is the counterpart and the trap. It works from whatever recording is at hand, which makes it easy to try and easy to commit to by accident, and a season built on a clone taken from a phone memo is expensive to re-voice.
2The clone travels further than the conversation did
One platform documents cloning alongside translation into thirty or more languages. The mechanics make that trivial; the voice being carried into thirty markets belongs to somebody, and only a release form makes it lawful.
No API quick start contains one, and this register does not expect it to. Recording the silence is the useful part, because a team reading only the mechanics will not be prompted to ask the question.
3One vendor documents cloning by ruling it out
A compliance checklist elsewhere in this register treats cloning an identifiable voice as a likeness question and names United States digital replica statutes. It is the only place here where a boundary appears at all.
It also leaves the boundary undrawn. A reference recording of a performer and a clone of an identifiable voice are not always easy to separate, and the checklist names the risk without saying where the line falls.
4Sources read for this entry
This page defines the term and logs what each vendor documents about it, each figure read on 2026-09-12. Related: Text to speech, Audio-driven video.