Kling AI: native audio, tied to a model version
Kling AI's guide puts native audio on VIDEO 3.0 and counts five languages for it. The capability therefore belongs to a model version, which is the thing a production has to pin down before a season is committed to it. As of 2026-09-12.
| What the documentation settles | What it leaves to a take |
|---|---|
| Native audio is documented on VIDEO 3.0, in five languages | Whether an earlier or cheaper model line carries the same behaviour |
| Voices are bound to elements, so a character keeps its voice between generations | What that voice will sound like before the first generation is made |
| Lip-sync is named as a capability of the model | What drives the mouth, and what happens with two speakers in frame |
Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.
1Attaching audio to a version is a scheduling fact
When a capability sits on a model line rather than on the account, a production is choosing a version and inheriting everything else that version does. A cheaper or faster line may render the same shot list without speaking, and the shot list will not say so. The check belongs at the start, not at the first dialogue scene.
It also means a version change is a dialogue change. Where a vendor moves capabilities between lines over time, the audio behaviour of a series can shift without anybody editing a prompt, which is an argument for recording the version alongside the output.
2Five is a count, and a count cannot be planned against
Five languages tells a reader the capability is multilingual and nothing more. Which five is the question a commissioning conversation asks, and it is not answered on the page. A production whose next market is one of the five learns nothing; a production whose next market is not learns nothing either.
The count sits oddly beside the element system, which is the most developed voice arrangement in this register. A vendor that has thought carefully about a voice persisting has published no list of the languages that voice can speak.
3The others that make sound while they make the picture
Seven more entries generate the audio in the same pass. What separates them is whether the vendor also says which languages, whose voice, and what happens to the mouth.
- LTX Studio — with the picture, and audio to video.
- MiniMax — with the picture.
- SceneMixer — with the picture, in the language set for the project.
- Sora 2 — with the picture.
- Veo — with the picture.
- Vidu — with the picture, speech available on its own.
- Wan 3.0 — with the picture, unless switched off.
- Audio sourceNative audio on VIDEO 3.0 in five languagesgenerated with the picture
- Per-character bindingVoices are bound to elements, so a character carries its voice between generationstied to the element system
- Lip-syncLip-sync is documentedstated without a mechanism
4Sources
Read from the model guide at kling.ai on 2026-09-12. The same column across every entry is on audio source; everything this vendor publishes about speech is on Kling AI. What counts as documented is on how read.