What Kling AI documents about speech
Kling AI (kling.ai), an AI creative studio, documents native audio on VIDEO 3.0 in five languages and binds a voice to an element, so a character keeps its voice across generations. As of 2026-09-12.
| Field | What the vendor documents |
|---|---|
| Audio source | Generated with the picture, VIDEO 3.0 |
| Languages | Five, not named |
| Voice source | Not documented by the vendor |
| Per-character binding | Voices bound to elements |
| Lip-sync | Documented, mechanism not described |
Inclusion rule. Fields are the same five for every model here; a field the vendor does not document is recorded as not documented rather than inferred from generated output. Order. Fixed field order, identical on every model page.
1Binding a voice to a character, not to a call
The element system is what makes this entry different from the others. A character defined once as an element carries its voice with it, which means the voice is a property of the character rather than something re-specified on every generation. For a serial production that is the difference between a cast and a series of coincidences.
What the guide does not give is a way to audition or pin a specific voice: the five languages are a count rather than a list, and no voice identifiers appear. A production can therefore rely on the voice staying the same without being able to say in advance what it will be.

2Lip-sync is claimed without a mechanism
Lip-sync appears in the documentation as a capability. Nothing states whether it is driven by the generated audio, by the text, or by a separate pass, and nothing states what happens with two speakers in frame. Those distinctions decide whether a dialogue scene can be shot in one generation, so they are recorded as unstated rather than assumed.
The same silence covers what happens when a supplied voice conflicts with the element's bound voice. Where a vendor documents a hierarchy, this site records it; Kling does not.
3This entry, one column at a time
Each of these stays inside a single field: what Kling AI puts there, what the wording settles, and what it leaves for a take to answer.
- Audio source — with the picture, on video 3.0.
- Voice source — nothing documented as an input.
- Per-character binding — on the element.
- Languages — a count, no names.
- Lip-sync — named, with nothing driving it.
4Read against another entry
Each of these puts Kling AI beside one other entry on a column where the two land at opposite grades of answer.
- SceneMixer and Kling AI — on languages.
- Kling AI and Synthesia — on lip-sync.
5Where this model sits against the rest
| Model | Audio source |
|---|---|
| D-ID | A text script read aloud, or an audio url supplied |
| Hedra | Supplied; audio is a required input |
| HeyGen | A text to speech endpoint, apart from the picture |
| Kling AI | Generated with the picture, VIDEO 3.0 |
| LTX Studio | Generated jointly with the picture, plus audio-to-video |
| Luma Ray | Not documented; the video reference does not mention sound |
| MiniMax | Native speech, generated with the picture |
| PixVerse | Supplied to a speech endpoint, or read from text there |
| Runway | Not documented where voices are defined |
| SceneMixer | Generated with the picture, in the language set for the project |
| Sora 2 | Generated with the picture; video and audio are both listed as output |
| sync-3 | Supplied, or read from text on the call |
| Synthesia | Synthesised from the script, or an uploaded recording |
| Veo | Generated with the video rather than added afterwards |
| Vidu | Q3 generates speech, effects and music natively, on by default |
| Wan 3.0 | On by default; false returns a file with no audio track |
| Wan2.2-S2V | Supplied; the card describes the model as audio-driven |
Inclusion rule. Models whose vendor documents something about generated speech. A model with nothing published on the point is still listed, with the cell marked hollow. Order. Alphabetical by model name.
| Model | Languages |
|---|---|
| D-ID | A language field beside the voice; no list on the reference |
| Hedra | Full multi-language support claimed, none named |
| HeyGen | Thirty or more, counted on the translation endpoint |
| Kling AI | Five, not named |
| LTX Studio | Not documented |
| Luma Ray | Not documented |
| MiniMax | Not documented by the vendor |
| PixVerse | Multiple, with speech, singing and advertisements named as types |
| Runway | Not documented |
| SceneMixer | 15, named, with Cantonese a 16th for dialogue only |
| Sora 2 | Not documented |
| sync-3 | Ninety-five or more, counted rather than named |
| Synthesia | A catalogue, each voice carrying a language name and code |
| Veo | Not documented |
| Vidu | Not documented |
| Wan 3.0 | Not documented |
| Wan2.2-S2V | Not documented |
Inclusion rule. Models whose vendor documents something about generated speech. A model with nothing published on the point is still listed, with the cell marked hollow. Order. Alphabetical by model name.
| Model | Voice source |
|---|---|
| D-ID | A voice id from one of five named speech providers |
| Hedra | An uploaded track, or a voice id from the voices endpoint |
| HeyGen | A clone, instant from one recording or professional from 20 minutes or more |
| Kling AI | Not documented by the vendor |
| LTX Studio | Not documented |
| Luma Ray | Not documented |
| MiniMax | Reference audio, 15 s total across 3 clips |
| PixVerse | Built-in voices, or a custom voice from a supplied sample |
| Runway | An audio sample of 10 seconds to 5 minutes, or a description in words |
| SceneMixer | Preset voice library, or the user's own recordings |
| Sora 2 | Not documented |
| sync-3 | Not documented on the model page |
| Synthesia | The catalogue, or a cloned voice with a language list of its own |
| Veo | Not documented |
| Vidu | Not documented |
| Wan 3.0 | Reference audio, 15 seconds in total, WAV or MP3 up to 15 MB |
| Wan2.2-S2V | Whatever track is handed over; no catalogue |
Inclusion rule. Models whose vendor documents something about generated speech. A model with nothing published on the point is still listed, with the cell marked hollow. Order. Alphabetical by model name.
6Sources
Taken from the model guide at kling.ai on 2026-09-12. Every entry against the same five fields sits on the speech controls page; the rule behind each field is on how read. The entries either side of this one: LTX Studio, Luma Ray.
- Audio sourceNative audio on VIDEO 3.0 in five languagesgenerated with the picture
- Per-character bindingVoices are bound to elements, so a character carries its voice between generationstied to the element system
- Lip-syncLip-sync is documentedstated without a mechanism