Kling AI: lip-sync as a capability, and nothing behind it
Lip-sync appears in the documentation as something the model does. Nothing states whether it follows the generated audio, the text, or a separate pass, and nothing addresses two people talking in one shot. As of 2026-09-12.
| What the documentation settles | What it leaves to a take |
|---|---|
| Lip-sync is named as a capability of the model | What drives the mouth: the audio, the text, or a later pass |
| Native audio is documented on VIDEO 3.0, in five languages | Whether alignment is equally good in all five |
| Voices are bound to elements, so a character carries its voice between generations | What happens if a bound voice and a supplied one disagree |
Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.
1Three mechanisms fail in visibly different ways
Generating sound and picture together, animating to a supplied track, and repairing a mouth afterwards produce different artefacts: soft or approximate movement, tight movement with a mismatched face, and a mouth that looks pasted on. A vendor naming only the feature has not told a production which of those to expect.
That is why this column grades a named driver above a named feature. The grade is not about confidence in the product; it is about how much a reader can plan from the sentence.
2The two-speaker case is the one that decides coverage
Whether a dialogue scene can be shot in one generation depends entirely on what happens when two faces are in frame and one is talking. Nothing here addresses it, so the safe assumption is alternating singles, which doubles the number of generations a conversation costs.
The element system makes that assumption cheaper to live with, because the voices stay consistent across the extra shots. Consistency across singles is a partial answer to a problem the documentation does not acknowledge.
3The others that name the feature and stop
Three more entries name lip-sync without naming a driver. Between them the claim sits in a translation sentence, in an endpoint name, and in a warning about long speeches.
- HeyGen — named, alongside translation.
- PixVerse — named as the endpoint's purpose.
- Sora 2 — named only where it fails.
- Lip-syncLip-sync is documentedstated without a mechanism
- Audio sourceNative audio on VIDEO 3.0 in five languagesgenerated with the picture
- Per-character bindingVoices are bound to elements, so a character carries its voice between generationstied to the element system
4Sources
Read from the model guide at kling.ai on 2026-09-12. The same column across every entry is on lip-sync; everything this vendor publishes about speech is on Kling AI. What counts as documented is on how read.