Sentioscope

Speech and voice controls, as each vendor documents them

Per-character binding: does a voice persist

Per-character binding records whether a voice is a property of a character that survives between generations, or something re-established each time. Four entries let a voice exist as a stored object. Five put it on the request. The remaining eight publish nothing here. As of 2026-09-12.

A voice that belongs to the character, against one supplied per callBoth arrangements give a consistent voice when everything goes right, and product copy describes them in the same words. What separates them is how a long series fails when one step is missed.Bound to the characterSupplied on the callConfiguredOnce, when the character is definedOnce per shot, every shotIf a step is missedNothing to missA different person, unannouncedChances to go wrongFlat across a seasonGrows with the episode countChoosing the voiceNot documented either wayWhatever the supplied clip producedSame promise, different failure at series length
Fig. 1 The field records where the voice lives, because that is what decides whether consistency survives a missed step at episode thirty.
Controls documented, model by modelFilled where the model documents that control, hollow where nothing is published about it.Controls documented, model by modelWhere the voice livesWhere the voicelivesThe published wording behind the cellThe publishedwording behind th…D-IDD-ID — Where the voice lives: On the requestD-ID — The published wording behind the cell: A voice id selected from the list of available voicesHedraHedra — Where the voice lives: Not documented by the vendorHedra — The published wording behind the cell: Audio or a voice id travels with each callHeyGenHeyGen — Where the voice lives: On a stored cloneHeyGen — The published wording behind the cell: Pass the clone's voice id in the requestKling AIKling AI — Where the voice lives: On the elementKling AI — The published wording behind the cell: A character carries its voice between generationsLTX StudioLTX Studio — Where the voice lives: On the character ElementLTX Studio — The published wording behind the cell: Voices are attached to ElementsLuma RayLuma Ray — Where the voice lives: Not documented by the vendorLuma Ray — The published wording behind the cell: No voice is described, so none can be attachedMiniMaxMiniMax — Where the voice lives: On the callMiniMax — The published wording behind the cell: A reference clip, capped across three clips, supplied each timePixVersePixVerse — Where the voice lives: On the requestPixVerse — The published wording behind the cell: The speaker id is passed on the generation requestRunwayRunway — Where the voice lives: On a stored voiceRunway — The published wording behind the cell: A voice becomes ready with a preview, then is used by idSceneMixerSceneMixer — Where the voice lives: Not documented by the vendorSceneMixer — The published wording behind the cell: Voice routes are described; persistence between shots is notSora 2Sora 2 — Where the voice lives: On the promptSora 2 — The published wording behind the cell: Speakers labelled consistently, with turns alternatedsync-3sync-3 — Where the voice lives: Not documented by the vendorsync-3 — The published wording behind the cell: Nothing published about a voice persisting between callsSynthesiaSynthesia — Where the voice lives: Not documented by the vendorSynthesia — The published wording behind the cell: A voice id exists; its pairing with an avatar is unstatedVeoVeo — Where the voice lives: Not documented by the vendorVeo — The published wording behind the cell: Nothing published about a voice surviving a generationViduVidu — Where the voice lives: Not documented by the vendorVidu — The published wording behind the cell: Audio is documented as an output of each callWan 3.0Wan 3.0 — Where the voice lives: On the callWan 3.0 — The published wording behind the cell: Reference audio is supplied per generationWan2.2-S2VWan2.2-S2V — Where the voice lives: Not documented by the vendorWan2.2-S2V — The published wording behind the cell: The voice is whoever made the recording
Fig. 2 Filled where the model documents that control, hollow where nothing is published about it.
Models documenting each controlHow many models document each control. A hollow column is a statement about documentation, not capability.Models documenting each controlWhere the voice lives9 of 17The published wording behind the ce…The published wording behind the cell16 of 17
Fig. 3 How many models document each control. A hollow column is a statement about documentation, not capability.
Whether a voice belongs to the character or to the call. Recorded 2026-09-12.
ModelWhere the voice livesThe published wording behind the cell
D-IDOn the requestA voice id selected from the list of available voices
HedraNot documented by the vendorAudio or a voice id travels with each call
HeyGenOn a stored clonePass the clone's voice id in the request
Kling AIOn the elementA character carries its voice between generations
LTX StudioOn the character ElementVoices are attached to Elements
Luma RayNot documented by the vendorNo voice is described, so none can be attached
MiniMaxOn the callA reference clip, capped across three clips, supplied each time
PixVerseOn the requestThe speaker id is passed on the generation request
RunwayOn a stored voiceA voice becomes ready with a preview, then is used by id
SceneMixerNot documented by the vendorVoice routes are described; persistence between shots is not
Sora 2On the promptSpeakers labelled consistently, with turns alternated
sync-3Not documented by the vendorNothing published about a voice persisting between calls
SynthesiaNot documented by the vendorA voice id exists; its pairing with an avatar is unstated
VeoNot documented by the vendorNothing published about a voice surviving a generation
ViduNot documented by the vendorAudio is documented as an output of each call
Wan 3.0On the callReference audio is supplied per generation
Wan2.2-S2VNot documented by the vendorThe voice is whoever made the recording

Inclusion rule. Models whose documentation says whether a voice attaches to a reusable character. Keeping a voice by re-supplying the same clip counts, and is recorded as living on the call. Order. Alphabetical by model name.

1Set once, or remembered by whoever presses the button

Both arrangements produce a consistent voice on a good day, which is why product copy describes them in the same words. They fail differently. A voice bound to an element is wrong once if it is wrong at all, and correcting it corrects every shot after it. A voice re-supplied per call is one more step in a repeated process, and every repeated step is a place where a season drifts.

The field therefore records where the voice lives rather than whether consistency is possible. Both routes reach the same destination; only one of them survives a tired operator at two in the morning forgetting to attach anything.

2Forty episodes is where the difference becomes visible

On a single scene the distinction is academic, and that is how it is usually demonstrated. At series length the arithmetic changes: a bound voice is configured once for a run, while a per-call reference is configured once per shot, so the number of chances to get it wrong grows with the length of the show rather than staying flat.

That is also why the clip belongs beside the character in whatever the production uses as an asset store. A shot regenerated months later from a different sample gives a different person, and nothing in the returned file announces the substitution.

3A bound voice is still not a voice anyone chose

Where a vendor documents binding, it documents persistence and not selection. The element system that keeps a voice stable does not come with a way to audition what that voice will be, and the language counts published beside it are counts rather than lists of identifiers.

So a production can rely on a character sounding the same in episode forty as in episode one without being able to say in advance what it sounds like. Where nothing at all is published, the cell stays blank instead of borrowing an answer from how the product appears to behave.

4One entry at a time on this column

A cell gets a page of its own where the vendor says something specific in it, or where its silence is unusual among the entries answering the same way. The remaining cells are left in the table above, because a page repeating one short phrase would be worse than a row carrying it.

Entries that let a voice exist as an object before any shot does:

Entries that re-establish the voice on every generation:

Entries that publish nothing about a voice outlasting one call:

  • Hedra — not documented by the vendor.
  • SceneMixer — not documented by the vendor.
  • Synthesia — not documented by the vendor.
  • Per-character binding
    Voices are bound to elements, so a character carries its voice between generationstied to the element systemKling AI, model guide / recorded 2026-09-12
  • Per-character binding
    Voices are attached to character Elementstied to the element systemLTX Studio / recorded 2026-09-12
  • Per-character binding
    Not documented by the vendor (as of 2026-09-22)no public descriptionVidu, API reference / recorded 2026-09-12
  • Per-character binding
    Not documented by the vendor (as of 2026-09-22)no public descriptionSceneMixer, compliance guide / recorded 2026-09-12
  • Voice source
    Reference audio is capped at 15 seconds in total across at most 3 clipsa hard published limitMiniMax, video generation guide / recorded 2026-09-12
  • Per-character binding
    A voice is created asynchronously, reaches a ready state with a preview, and is then referred to by its own ida stored voice objectRunway, custom voices reference / recorded 2026-09-22
  • Per-character binding
    The voice is a voice id selected from the list of available voices, with an optional language field beside ita voice id on every requestD-ID, create a talk reference / recorded 2026-09-22
  • Per-character binding
    The chosen speaker id is passed as the lip-sync speaker on the generation requesta speaker id per callPixVerse, speech and lip sync guide / recorded 2026-09-22
  • Two characters in frame
    For multi-character scenes the guide asks for speakers labelled consistently and turns alternated, so each line lands on the right characteraddressed rather than skippedOpenAI, Sora 2 prompting guide / recorded 2026-09-22

5Sources

Each cell is read from the vendor page it links to, checked 2026-09-12. The fields sit side by side on the speech table, and what counts as documented is set out on how read. The other fields: Languages, Lip-sync. All of them: the field notes.