Per-character binding: does a voice persist
Per-character binding records whether a voice is a property of a character that survives between generations, or something re-established each time. Four entries let a voice exist as a stored object. Five put it on the request. The remaining eight publish nothing here. As of 2026-09-12.
| Model | Where the voice lives | The published wording behind the cell |
|---|---|---|
| D-ID | On the request | A voice id selected from the list of available voices |
| Hedra | Not documented by the vendor | Audio or a voice id travels with each call |
| HeyGen | On a stored clone | Pass the clone's voice id in the request |
| Kling AI | On the element | A character carries its voice between generations |
| LTX Studio | On the character Element | Voices are attached to Elements |
| Luma Ray | Not documented by the vendor | No voice is described, so none can be attached |
| MiniMax | On the call | A reference clip, capped across three clips, supplied each time |
| PixVerse | On the request | The speaker id is passed on the generation request |
| Runway | On a stored voice | A voice becomes ready with a preview, then is used by id |
| SceneMixer | Not documented by the vendor | Voice routes are described; persistence between shots is not |
| Sora 2 | On the prompt | Speakers labelled consistently, with turns alternated |
| sync-3 | Not documented by the vendor | Nothing published about a voice persisting between calls |
| Synthesia | Not documented by the vendor | A voice id exists; its pairing with an avatar is unstated |
| Veo | Not documented by the vendor | Nothing published about a voice surviving a generation |
| Vidu | Not documented by the vendor | Audio is documented as an output of each call |
| Wan 3.0 | On the call | Reference audio is supplied per generation |
| Wan2.2-S2V | Not documented by the vendor | The voice is whoever made the recording |
Inclusion rule. Models whose documentation says whether a voice attaches to a reusable character. Keeping a voice by re-supplying the same clip counts, and is recorded as living on the call. Order. Alphabetical by model name.
1Set once, or remembered by whoever presses the button
Both arrangements produce a consistent voice on a good day, which is why product copy describes them in the same words. They fail differently. A voice bound to an element is wrong once if it is wrong at all, and correcting it corrects every shot after it. A voice re-supplied per call is one more step in a repeated process, and every repeated step is a place where a season drifts.
The field therefore records where the voice lives rather than whether consistency is possible. Both routes reach the same destination; only one of them survives a tired operator at two in the morning forgetting to attach anything.
2Forty episodes is where the difference becomes visible
On a single scene the distinction is academic, and that is how it is usually demonstrated. At series length the arithmetic changes: a bound voice is configured once for a run, while a per-call reference is configured once per shot, so the number of chances to get it wrong grows with the length of the show rather than staying flat.
That is also why the clip belongs beside the character in whatever the production uses as an asset store. A shot regenerated months later from a different sample gives a different person, and nothing in the returned file announces the substitution.
3A bound voice is still not a voice anyone chose
Where a vendor documents binding, it documents persistence and not selection. The element system that keeps a voice stable does not come with a way to audition what that voice will be, and the language counts published beside it are counts rather than lists of identifiers.
So a production can rely on a character sounding the same in episode forty as in episode one without being able to say in advance what it sounds like. Where nothing at all is published, the cell stays blank instead of borrowing an answer from how the product appears to behave.
4One entry at a time on this column
A cell gets a page of its own where the vendor says something specific in it, or where its silence is unusual among the entries answering the same way. The remaining cells are left in the table above, because a page repeating one short phrase would be worse than a row carrying it.
Entries that let a voice exist as an object before any shot does:
- HeyGen — on a stored clone.
- Kling AI — on the element.
- LTX Studio — on the character element.
- Runway — on a stored voice.
Entries that re-establish the voice on every generation:
- D-ID — on the request.
- MiniMax — on the call.
- PixVerse — on the request.
- Sora 2 — on the prompt.
- Wan 3.0 — on the call.
Entries that publish nothing about a voice outlasting one call:
- Hedra — not documented by the vendor.
- SceneMixer — not documented by the vendor.
- Synthesia — not documented by the vendor.
- Per-character bindingVoices are bound to elements, so a character carries its voice between generationstied to the element system
- Per-character bindingVoices are attached to character Elementstied to the element system
- Per-character bindingNot documented by the vendor (as of 2026-09-22)no public description
- Per-character bindingNot documented by the vendor (as of 2026-09-22)no public description
- Voice sourceReference audio is capped at 15 seconds in total across at most 3 clipsa hard published limit
- Per-character bindingA voice is created asynchronously, reaches a ready state with a preview, and is then referred to by its own ida stored voice object
- Per-character bindingThe voice is a voice id selected from the list of available voices, with an optional language field beside ita voice id on every request
- Per-character bindingThe chosen speaker id is passed as the lip-sync speaker on the generation requesta speaker id per call
- Two characters in frameFor multi-character scenes the guide asks for speakers labelled consistently and turns alternated, so each line lands on the right characteraddressed rather than skipped
5Sources
Each cell is read from the vendor page it links to, checked 2026-09-12. The fields sit side by side on the speech table, and what counts as documented is set out on how read. The other fields: Languages, Lip-sync. All of them: the field notes.