How a character keeps the same voice
Four of the 17 models let a voice exist as something stored and reused: bound to an element, attached to a character, or created as a named object with a preview. Five re-establish it on each call from a clip, a speaker id or a provider voice. The remainder publish nothing. As of 2026-09-12.
| Model | What is documented | Detail |
|---|---|---|
| Kling AI | Bound to an element | Survives between generations |
| LTX Studio | Attached to a character Element | Survives between generations |
| HeyGen | A cloned voice, reused by id | Enrolled once, from one recording or twenty minutes |
| Runway | A stored voice with its own id | Ready with a preview before any shot exists |
| D-ID | A voice id on each request | Chosen from a provider catalogue |
| MiniMax | Reference audio per call | 15 seconds total across 3 clips |
| PixVerse | A speaker id per generation | Built-in, or custom from a supplied sample |
| Sora 2 | A label inside the prompt | Routes a line to a face, not an identity |
| Wan 3.0 | Reference audio per call | 15 seconds in total, WAV or MP3 |
| Hedra | Not documented | Audio or a voice id travels with the call |
| Luma Ray | Not documented | No voice is described at all |
| SceneMixer | Not documented | A preset library exists; persistence unstated |
| sync-3 | Not documented | Nothing published about a voice lasting |
| Synthesia | Not documented | A voice id exists; the pairing is unstated |
| Veo | Not documented | Nothing published about voices |
| Vidu | Not documented | No statement about voice at all |
| Wan2.2-S2V | Not documented | The voice is whoever recorded the track |
Inclusion rule. Every entry in the register. A model that documents no persistence mechanism is recorded as not documented rather than as lacking one. Order. Entries documenting a stored voice first, then per-call routes, then the rest alphabetically.
1Binding and re-supplying both produce consistency, and they fail differently
A bound voice is set once. It is wrong once if it is wrong at all, and it cannot drift because nothing is re-specified. A per-call reference is a step in every generation, and every repeated step is somewhere a long series can come apart: a clip forgotten, a different clip used, a slightly different take.
For forty episodes that difference compounds in one direction only. Neither vendor frames it as a consistency feature; the property falls out of how the API is shaped.
2Fifteen seconds, and no other figure in this column
Among the entries read first, one number: 15 seconds in total, across at most three clips. Every other cell in this column holds a mechanism with no limits attached, or holds nothing. A figure carries weight here because it shapes the work: a voice gets established from a very short sample, and no documented route exists for storing it.
A cap that small also constrains casting. Three clips of five seconds is not much material to characterise a voice from, and the documentation says nothing about what happens when the sample is unrepresentative.
3A library without a persistence rule leaves the question half-answered
One entry names a preset library and a route for supplying recordings, which settles where voices come from and not how one stays with a character across episodes. Those are separate questions and only the first is addressed.
This register keeps them as separate fields for that reason. Merging them would let a documented library imply a persistence guarantee that nobody has published.
- Per-character bindingVoices are bound to elements, so a character carries its voice between generationstied to the element system
- Per-character bindingVoices are attached to character Elementstied to the element system
- Voice sourceReference audio is capped at 15 seconds in total across at most 3 clipsa hard published limit
- Per-character bindingNot documented by the vendor (as of 2026-09-22)no public description
- Per-character bindingNot documented by the vendor (as of 2026-09-22)no public description
- Per-character bindingA voice is created asynchronously, reaches a ready state with a preview, and is then referred to by its own ida stored voice object
- Voice sourceA voice clone is instant from one recording or professional from twenty minutes or more, and is then passed as a voice idtwo enrolment routes with a threshold
- Per-character bindingThe voice is a voice id selected from the list of available voices, with an optional language field beside ita voice id on every request
- Per-character bindingThe chosen speaker id is passed as the lip-sync speaker on the generation requesta speaker id per call
- Voice sourceReference audio is WAV or MP3, one to fifteen seconds a clip, fifteen seconds in total and no more than 15 MBa published cap
4Sources
Every line is read from the vendor documentation linked on the speech controls page, checked 2026-09-12. What counts as documented is on how read. Nearby: Lip-sync, Voice sources, Sound by default.