Sentioscope

Speech and voice controls, as each vendor documents them

Kling AI: lip-sync as a capability, and nothing behind it

Lip-sync appears in the documentation as something the model does. Nothing states whether it follows the generated audio, the text, or a separate pass, and nothing addresses two people talking in one shot. As of 2026-09-12.

Three mechanisms, three visibly different failuresGenerating sound and picture together, animating to a supplied track, and repairing a mouth afterwards each fail in their own way. A vendor naming only the feature has not said which artefact to expect.What a driver statement would settleMade togetherMovement can be soft or approximate, because nothing externalsets it.Animated to audioTight movement, with a face that may not match the rest of theshot.Repaired afterwardsA mouth that can look pasted onto footage it did not come from.Named only as a featureWhere this entry sits: no way to tell which of the threeapplies.And what a capability claim leaves open
Fig. 1 The grade this column gives is about how much a reader can plan from the sentence, not about confidence in the product.
Kling AI on lip-sync, statement by statement. Read from the vendor's model guide on 2026-09-12.
What the documentation settlesWhat it leaves to a take
Lip-sync is named as a capability of the modelWhat drives the mouth: the audio, the text, or a later pass
Native audio is documented on VIDEO 3.0, in five languagesWhether alignment is equally good in all five
Voices are bound to elements, so a character carries its voice between generationsWhat happens if a bound voice and a supplied one disagree

Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.

1Three mechanisms fail in visibly different ways

Generating sound and picture together, animating to a supplied track, and repairing a mouth afterwards produce different artefacts: soft or approximate movement, tight movement with a mismatched face, and a mouth that looks pasted on. A vendor naming only the feature has not told a production which of those to expect.

That is why this column grades a named driver above a named feature. The grade is not about confidence in the product; it is about how much a reader can plan from the sentence.

2The two-speaker case is the one that decides coverage

Whether a dialogue scene can be shot in one generation depends entirely on what happens when two faces are in frame and one is talking. Nothing here addresses it, so the safe assumption is alternating singles, which doubles the number of generations a conversation costs.

The element system makes that assumption cheaper to live with, because the voices stay consistent across the extra shots. Consistency across singles is a partial answer to a problem the documentation does not acknowledge.

3The others that name the feature and stop

Three more entries name lip-sync without naming a driver. Between them the claim sits in a translation sentence, in an endpoint name, and in a warning about long speeches.

  • HeyGen — named, alongside translation.
  • PixVerse — named as the endpoint's purpose.
  • Sora 2 — named only where it fails.
  • Lip-sync
    Lip-sync is documentedstated without a mechanismKling AI, model guide / recorded 2026-09-12
  • Audio source
    Native audio on VIDEO 3.0 in five languagesgenerated with the pictureKling AI, model guide / recorded 2026-09-12
  • Per-character binding
    Voices are bound to elements, so a character carries its voice between generationstied to the element systemKling AI, model guide / recorded 2026-09-12

4Sources

Read from the model guide at kling.ai on 2026-09-12. The same column across every entry is on lip-sync; everything this vendor publishes about speech is on Kling AI. What counts as documented is on how read.