Sentioscope

Speech and voice controls, as each vendor documents them

Kling AI: the voice is a property of the element

Voices are bound to elements, and an element is the character. That makes the voice a property of the person rather than of the request, which is the only arrangement here that needs no per-shot discipline at all. As of 2026-09-12.

Nothing to pass means nothing to pass wronglyWhere the voice lives on the request, a production can supply the wrong one and nothing objects. Where it lives on the element there is nothing to supply, so the commonest continuity error is unavailable rather than guarded against.Voice on the elementVoice on the requestWhat a shot has to carryNothingA clip or an identifier, every timeCommonest failureNone availableThe wrong voice, silentlyWho chose the voiceNobody; the first generation didWhoever picked from a catalogueAuditioningNot describedSometimes, with a previewConsistency and casting are traded against each other
Fig. 1 The system that guarantees consistency is the same one that withholds the casting decision, and the two cannot be separated.
Kling AI on per-character binding, statement by statement. Read from the vendor's model guide on 2026-09-12.
What the documentation settlesWhat it leaves to a take
Voices are bound to elements, so a character carries its voice between generationsWhat that voice is, since nothing describes where it came from
Native audio is documented on VIDEO 3.0, in five languagesWhether an element's voice follows it across those languages
Lip-sync is named as a capability of the modelWhether a bound voice affects which mouth moves in a two-shot

Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.

1Putting the voice on the character removes a whole class of mistake

Everywhere the voice lives on the request, a production can pass the wrong one and nothing will complain. Here there is nothing to pass. The element carries its voice, so the commonest continuity error in generated dialogue is unavailable rather than guarded against.

That is worth more than it sounds over forty episodes, because the failure it prevents is silent. A voice that drifts halfway through a season is noticed by an audience long before it is noticed in a log.

2The cost is that nobody chose the voice

An element's voice is not described as selectable, auditionable or replaceable. The system that guarantees consistency is the same system that withholds the casting decision, and the two cannot be separated from what the guide publishes.

For a production that makes the first generation of each character consequential. Whatever comes back becomes the character for the rest of the series, and the guide describes no way to ask for a second opinion.

3The others where a voice exists before any shot does

Three more entries let a voice outlast a call. One other binds it to a character object; two create a voice object instead, which can be auditioned but has to be attached by hand.

  • Per-character binding
    Voices are bound to elements, so a character carries its voice between generationstied to the element systemKling AI, model guide / recorded 2026-09-12
  • Audio source
    Native audio on VIDEO 3.0 in five languagesgenerated with the pictureKling AI, model guide / recorded 2026-09-12
  • Lip-sync
    Lip-sync is documentedstated without a mechanismKling AI, model guide / recorded 2026-09-12

4Sources

Read from the model guide at kling.ai on 2026-09-12. The same column across every entry is on per-character binding; everything this vendor publishes about speech is on Kling AI. What counts as documented is on how read.