Kling AI: the voice is a property of the element
Voices are bound to elements, and an element is the character. That makes the voice a property of the person rather than of the request, which is the only arrangement here that needs no per-shot discipline at all. As of 2026-09-12.
| What the documentation settles | What it leaves to a take |
|---|---|
| Voices are bound to elements, so a character carries its voice between generations | What that voice is, since nothing describes where it came from |
| Native audio is documented on VIDEO 3.0, in five languages | Whether an element's voice follows it across those languages |
| Lip-sync is named as a capability of the model | Whether a bound voice affects which mouth moves in a two-shot |
Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.
1Putting the voice on the character removes a whole class of mistake
Everywhere the voice lives on the request, a production can pass the wrong one and nothing will complain. Here there is nothing to pass. The element carries its voice, so the commonest continuity error in generated dialogue is unavailable rather than guarded against.
That is worth more than it sounds over forty episodes, because the failure it prevents is silent. A voice that drifts halfway through a season is noticed by an audience long before it is noticed in a log.
2The cost is that nobody chose the voice
An element's voice is not described as selectable, auditionable or replaceable. The system that guarantees consistency is the same system that withholds the casting decision, and the two cannot be separated from what the guide publishes.
For a production that makes the first generation of each character consequential. Whatever comes back becomes the character for the rest of the series, and the guide describes no way to ask for a second opinion.
3The others where a voice exists before any shot does
Three more entries let a voice outlast a call. One other binds it to a character object; two create a voice object instead, which can be auditioned but has to be attached by hand.
- HeyGen — on a stored clone.
- LTX Studio — on the character element.
- Runway — on a stored voice.
- Per-character bindingVoices are bound to elements, so a character carries its voice between generationstied to the element system
- Audio sourceNative audio on VIDEO 3.0 in five languagesgenerated with the picture
- Lip-syncLip-sync is documentedstated without a mechanism
4Sources
Read from the model guide at kling.ai on 2026-09-12. The same column across every entry is on per-character binding; everything this vendor publishes about speech is on Kling AI. What counts as documented is on how read.