Sentioscope

Speech and voice controls, as each vendor documents them

Kling AI: native audio, tied to a model version

Kling AI's guide puts native audio on VIDEO 3.0 and counts five languages for it. The capability therefore belongs to a model version, which is the thing a production has to pin down before a season is committed to it. As of 2026-09-12.

Audio that belongs to a model version, not to an accountNative audio is documented on VIDEO 3.0 and counted in five languages, and voices bind to elements. Each of those sits at a different level, so a production choosing a cheaper line can lose the first without touching the third.Where each published claim actually attachesModel versionNative audio is documented on VIDEO 3.0; other lines are notaddressed.Language countFive languages are counted for that audio, and none of them isnamed.ElementA voice binds to a character element and travels with it betweengenerations.A version change is a dialogue change
Fig. 1 Recording the version alongside the output is the habit this arrangement quietly asks for.
Kling AI on audio source, statement by statement. Read from the vendor's model guide on 2026-09-12.
What the documentation settlesWhat it leaves to a take
Native audio is documented on VIDEO 3.0, in five languagesWhether an earlier or cheaper model line carries the same behaviour
Voices are bound to elements, so a character keeps its voice between generationsWhat that voice will sound like before the first generation is made
Lip-sync is named as a capability of the modelWhat drives the mouth, and what happens with two speakers in frame

Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.

1Attaching audio to a version is a scheduling fact

When a capability sits on a model line rather than on the account, a production is choosing a version and inheriting everything else that version does. A cheaper or faster line may render the same shot list without speaking, and the shot list will not say so. The check belongs at the start, not at the first dialogue scene.

It also means a version change is a dialogue change. Where a vendor moves capabilities between lines over time, the audio behaviour of a series can shift without anybody editing a prompt, which is an argument for recording the version alongside the output.

2Five is a count, and a count cannot be planned against

Five languages tells a reader the capability is multilingual and nothing more. Which five is the question a commissioning conversation asks, and it is not answered on the page. A production whose next market is one of the five learns nothing; a production whose next market is not learns nothing either.

The count sits oddly beside the element system, which is the most developed voice arrangement in this register. A vendor that has thought carefully about a voice persisting has published no list of the languages that voice can speak.

3The others that make sound while they make the picture

Seven more entries generate the audio in the same pass. What separates them is whether the vendor also says which languages, whose voice, and what happens to the mouth.

  • LTX Studio — with the picture, and audio to video.
  • MiniMax — with the picture.
  • SceneMixer — with the picture, in the language set for the project.
  • Sora 2 — with the picture.
  • Veo — with the picture.
  • Vidu — with the picture, speech available on its own.
  • Wan 3.0 — with the picture, unless switched off.
  • Audio source
    Native audio on VIDEO 3.0 in five languagesgenerated with the pictureKling AI, model guide / recorded 2026-09-12
  • Per-character binding
    Voices are bound to elements, so a character carries its voice between generationstied to the element systemKling AI, model guide / recorded 2026-09-12
  • Lip-sync
    Lip-sync is documentedstated without a mechanismKling AI, model guide / recorded 2026-09-12

4Sources

Read from the model guide at kling.ai on 2026-09-12. The same column across every entry is on audio source; everything this vendor publishes about speech is on Kling AI. What counts as documented is on how read.