Sentioscope

Speech and voice controls, as each vendor documents them

How a character keeps the same voice

Four of the 17 models let a voice exist as something stored and reused: bound to an element, attached to a character, or created as a named object with a preview. Five re-establish it on each call from a clip, a speaker id or a provider voice. The remainder publish nothing. As of 2026-09-12.

Controls documented, model by modelFilled where the model documents that control, hollow where nothing is published about it.Controls documented, model by modelWhat is documentedKling AIKling AI — What is documented: Bound to an elementLTX StudioLTX Studio — What is documented: Attached to a character ElementHeyGenHeyGen — What is documented: A cloned voice, reused by idRunwayRunway — What is documented: A stored voice with its own idD-IDD-ID — What is documented: A voice id on each requestMiniMaxMiniMax — What is documented: Reference audio per callPixVersePixVerse — What is documented: A speaker id per generationSora 2Sora 2 — What is documented: A label inside the promptWan 3.0Wan 3.0 — What is documented: Reference audio per callHedraHedra — What is documented: Not documentedLuma RayLuma Ray — What is documented: Not documentedSceneMixerSceneMixer — What is documented: Not documentedsync-3sync-3 — What is documented: Not documentedSynthesiaSynthesia — What is documented: Not documentedVeoVeo — What is documented: Not documentedViduVidu — What is documented: Not documentedWan2.2-S2VWan2.2-S2V — What is documented: Not documented
Fig. 1 Filled where the model documents that control, hollow where nothing is published about it.
How each entry documents a voice staying attached to a character. Recorded 2026-09-12.
ModelWhat is documentedDetail
Kling AIBound to an elementSurvives between generations
LTX StudioAttached to a character ElementSurvives between generations
HeyGenA cloned voice, reused by idEnrolled once, from one recording or twenty minutes
RunwayA stored voice with its own idReady with a preview before any shot exists
D-IDA voice id on each requestChosen from a provider catalogue
MiniMaxReference audio per call15 seconds total across 3 clips
PixVerseA speaker id per generationBuilt-in, or custom from a supplied sample
Sora 2A label inside the promptRoutes a line to a face, not an identity
Wan 3.0Reference audio per call15 seconds in total, WAV or MP3
HedraNot documentedAudio or a voice id travels with the call
Luma RayNot documentedNo voice is described at all
SceneMixerNot documentedA preset library exists; persistence unstated
sync-3Not documentedNothing published about a voice lasting
SynthesiaNot documentedA voice id exists; the pairing is unstated
VeoNot documentedNothing published about voices
ViduNot documentedNo statement about voice at all
Wan2.2-S2VNot documentedThe voice is whoever recorded the track

Inclusion rule. Every entry in the register. A model that documents no persistence mechanism is recorded as not documented rather than as lacking one. Order. Entries documenting a stored voice first, then per-call routes, then the rest alphabetically.

1Binding and re-supplying both produce consistency, and they fail differently

A bound voice is set once. It is wrong once if it is wrong at all, and it cannot drift because nothing is re-specified. A per-call reference is a step in every generation, and every repeated step is somewhere a long series can come apart: a clip forgotten, a different clip used, a slightly different take.

For forty episodes that difference compounds in one direction only. Neither vendor frames it as a consistency feature; the property falls out of how the API is shaped.

2Fifteen seconds, and no other figure in this column

Among the entries read first, one number: 15 seconds in total, across at most three clips. Every other cell in this column holds a mechanism with no limits attached, or holds nothing. A figure carries weight here because it shapes the work: a voice gets established from a very short sample, and no documented route exists for storing it.

A cap that small also constrains casting. Three clips of five seconds is not much material to characterise a voice from, and the documentation says nothing about what happens when the sample is unrepresentative.

3A library without a persistence rule leaves the question half-answered

One entry names a preset library and a route for supplying recordings, which settles where voices come from and not how one stays with a character across episodes. Those are separate questions and only the first is addressed.

This register keeps them as separate fields for that reason. Merging them would let a documented library imply a persistence guarantee that nobody has published.

  • Per-character binding
    Voices are bound to elements, so a character carries its voice between generationstied to the element systemKling AI, model guide / recorded 2026-09-12
  • Per-character binding
    Voices are attached to character Elementstied to the element systemLTX Studio / recorded 2026-09-12
  • Voice source
    Reference audio is capped at 15 seconds in total across at most 3 clipsa hard published limitMiniMax, video generation guide / recorded 2026-09-12
  • Per-character binding
    Not documented by the vendor (as of 2026-09-22)no public descriptionSceneMixer, compliance guide / recorded 2026-09-12
  • Per-character binding
    Not documented by the vendor (as of 2026-09-22)no public descriptionVidu, API reference / recorded 2026-09-12
  • Per-character binding
    A voice is created asynchronously, reaches a ready state with a preview, and is then referred to by its own ida stored voice objectRunway, custom voices reference / recorded 2026-09-22
  • Voice source
    A voice clone is instant from one recording or professional from twenty minutes or more, and is then passed as a voice idtwo enrolment routes with a thresholdHeyGen, API quick start / recorded 2026-09-22
  • Per-character binding
    The voice is a voice id selected from the list of available voices, with an optional language field beside ita voice id on every requestD-ID, create a talk reference / recorded 2026-09-22
  • Per-character binding
    The chosen speaker id is passed as the lip-sync speaker on the generation requesta speaker id per callPixVerse, speech and lip sync guide / recorded 2026-09-22
  • Voice source
    Reference audio is WAV or MP3, one to fifteen seconds a clip, fifteen seconds in total and no more than 15 MBa published capAlibaba Cloud, Wan 3.0 API reference / recorded 2026-09-22

4Sources

Every line is read from the vendor documentation linked on the speech controls page, checked 2026-09-12. What counts as documented is on how read. Nearby: Lip-sync, Voice sources, Sound by default.