Voice source: what a production hands over
Voice source records what a production can hand over so speech comes out as a particular character. Ten of the 17 entries answer it, between them describing a catalogue, a clone enrolled once, a clip supplied on the call, and a voice asked for in words. As of 2026-09-12.
| Model | What can be supplied | The published wording behind the cell |
|---|---|---|
| D-ID | A voice id, or a recording by url | Five providers named, from Microsoft to Azure OpenAI |
| Hedra | A track, or a voice id from the endpoint | Audio uploaded, or speech generated inline |
| HeyGen | A recording to clone from | Instant from one recording, professional from twenty minutes |
| Kling AI | Nothing documented as an input | A voice arrives with the element, not from the user |
| LTX Studio | Nothing documented as an input | Voices are attached to character Elements |
| Luma Ray | Nothing documented as an input | Nothing published about voices anywhere on the page |
| MiniMax | A reference clip on the call | Fifteen seconds in total, across at most three clips |
| PixVerse | A sample, or a built-in voice | Built-in voices and custom voices from supplied audio |
| Runway | A 10-second to 5-minute sample, or a sentence | At most 10 MB, or a description of 20 characters |
| SceneMixer | A preset library, or the user's own recordings | Named in a compliance checklist rather than in product copy |
| Sora 2 | Nothing documented as an input | Speakers are labelled; their voices are not specified |
| sync-3 | Nothing documented as an input | Text is listed as an input, voices are not discussed |
| Synthesia | A catalogue voice, or a cloned one | Each voice carries a language, a gender and an id |
| Veo | Not documented by the vendor | Audio is documented; its origin is not addressed |
| Vidu | Nothing documented as an input | Audio is described as an output of the call |
| Wan 3.0 | Reference audio, 15 seconds in total | WAV or MP3, one to fifteen seconds a clip, 15 MB |
| Wan2.2-S2V | The audio track itself | An audio input with a reference image and a prompt |
Inclusion rule. Models whose documentation says what a production may hand over to steer the voice. A model that produces voices without accepting anything from the user is recorded as documenting no input. Order. Alphabetical by model name.
1A library and a clip are not the same offer
Choosing from a preset library and supplying a recording answer the same production question in opposite directions. A library is a closed set that can be auditioned before a season starts, so casting happens once and the choice is repeatable. A reference clip is an open set, so the voice is whatever the sample produced and the sample becomes an asset to be stored.
The register keeps both in one field because a reader is asking where the voice comes from, not which mechanism the vendor prefers. What the field cannot do is rank them: a library that does not contain the voice a scene needs is worse than a clip that does.
2Fifteen seconds is the figure that shapes the casting
One entry here publishes a number rather than a capability: reference audio capped at fifteen seconds in total, split across no more than three clips. That is enough to carry a timbre and not enough to carry a reading, so the clip should be chosen for the voice rather than for the best take of the line.
The cap also sits on the call, not on the project, so a scene with several speaking parts divides the same budget. Nothing in that documentation describes storing the result, which means the clip goes in again every time the character speaks.
3The field where consent turns up
The clearest sentence in this column came out of a compliance checklist rather than a feature page: preset voices or the user's own recordings, with cloning an identifiable voice described as a likeness question and US digital-replica statutes named beside it. Obligation produced the specificity that marketing had not.
That is worth noting as a pattern for the rest of the register. Where a vendor has a reason to write something down, the writing is precise; where the only reason is to sell a capability, the wording stops at the capability and this field stays empty.
4One entry at a time on this column
A cell gets a page of its own where the vendor says something specific in it, or where its silence is unusual among the entries answering the same way. The remaining cells are left in the table above, because a page repeating one short phrase would be worse than a row carrying it.
Entries that take a clip on the call that makes it:
- MiniMax — a reference clip on the call.
- Wan 3.0 — reference audio, 15 seconds in total.
- Wan2.2-S2V — the audio track itself.
Entries that build a voice once and address it afterwards by name:
Entries that offer a catalogue to choose from:
- D-ID — a voice id, or a recording by url.
- Hedra — a track, or a voice id from the endpoint.
- PixVerse — a sample, or a built-in voice.
- SceneMixer — a preset library, or the user's own recordings.
- Synthesia — a catalogue voice, or a cloned one.
Entries that publish nothing a production could hand over:
- Kling AI — nothing documented as an input.
- LTX Studio — nothing documented as an input.
- sync-3 — nothing documented as an input.
- Voice sourceThe guide tells users to use a preset voice library or their own recordingstwo routes named
- Voice sourceReference audio is capped at 15 seconds in total across at most 3 clipsa hard published limit
- What the vendor rules outCloning an identifiable voice is described as a likeness issue, with US digital-replica statutes named
- Voice sourceA custom voice can be built from an audio sample between 10 seconds and 5 minutes long and at most 10 MBa published sample window
- Voice sourceA voice can instead be asked for in words, with the description required to run at least 20 charactersa route that needs no recording
- Voice sourceReference audio is WAV or MP3, one to fifteen seconds a clip, fifteen seconds in total and no more than 15 MBa published cap
- Voice sourceA voice clone is instant from one recording or professional from twenty minutes or more, and is then passed as a voice idtwo enrolment routes with a threshold
- Voice sourceText to speech supports both built-in voices and custom voices created from user-provided sample audiocatalogue or supplied sample
- Voice sourceVoice cloning carries a language list of its own, running from Afrikaans to Zulu and described as the same list used for standalone cloninga cloned voice with its own reach
- Voice sourceSpeech can be generated inline instead of uploaded, by naming a voice id drawn from the voices endpointa catalogue behind an endpoint
5Sources
Each cell is read from the vendor page it links to, checked 2026-09-12. The fields sit side by side on the speech table, and what counts as documented is set out on how read. The other fields: Per-character binding, Languages. All of them: the field notes.