Sentioscope

Speech and voice controls, as each vendor documents them

Kling AI and Synthesia: a capability, or a mechanism

One entry names lip-sync as a capability of the model and describes nothing behind it. The other says lip sync and facial expressions are generated from the spoken content, and that closer framing performs better. As of 2026-09-22.

Naming a capability, against naming a mechanismThree mechanisms can put a moving mouth on screen and they fail visibly differently. A page naming only the feature has not said which artefact to expect, so a bad shot cannot be diagnosed.Kling AISynthesiaThe lip-sync claimA capability of the modelGenerated from the spoken contentA weak case namedNoneWide framing performs less wellTwo speakersNot addressedImplied by the framing conditionThe voiceBound to a character, source unstatedPublished with a code, pairing unstatedPersistence without selection, against selection without persistence
Fig. 1 The grade this register gives says nothing about which model aligns better; it says how much a reader can plan from the sentence.
Controls documented, model by modelFilled where the model documents that control, hollow where nothing is published about it.Controls documented, model by modelKling AISynthesiaWhere they partAudio sourceAudio source — Kling AI: With the picture, on VIDEO 3.0Audio source — Synthesia: From the script, or uploadedAudio source — Where they part: Different answersVoice sourceVoice source — Kling AI: Nothing documented as an inputVoice source — Synthesia: A catalogue voice, or a cloned oneVoice source — Where they part: Different answersPer-character bindingPer-character binding — Kling AI: On the elementPer-character binding — Synthesia: Not documented by the vendorPer-character binding — Where they part: Only Kling AI answersLanguagesLanguages — Kling AI: A count, no namesLanguages — Synthesia: A catalogue with codesLanguages — Where they part: Different answersLip-syncLip-sync — Kling AI: Named, with nothing driving itLip-sync — Synthesia: The spoken content, framed closeLip-sync — Where they part: Different answers
Fig. 2 Filled where the model documents that control, hollow where nothing is published about it.
Models documenting each controlHow many models document each control. A hollow column is a statement about documentation, not capability.Models documenting each controlKling AI4 of 5Synthesia4 of 5Where they part5 of 5
Fig. 3 How many models document each control. A hollow column is a statement about documentation, not capability.
Kling AI and Synthesia on the five fields this register keeps, with the gap on each. Recorded 2026-09-22.
FieldKling AISynthesiaWhere they part
Audio sourceWith the picture, on VIDEO 3.0From the script, or uploadedDifferent answers
Voice sourceNothing documented as an inputA catalogue voice, or a cloned oneDifferent answers
Per-character bindingOn the elementNot documented by the vendorOnly Kling AI answers
LanguagesA count, no namesA catalogue with codesDifferent answers
Lip-syncNamed, with nothing driving itThe spoken content, framed closeDifferent answers

Inclusion rule. Two entries are given a page together when at least one column puts them at opposite grades of answer. Pairs that agree on every column, or that are both blank throughout, do not get a page. Order. Fixed field order, identical on every side-by-side page.

1A capability claim and a driver statement are not the same grade

Three mechanisms can put a moving mouth on screen, and they fail in visibly different ways. A page naming only the feature has not said which artefact to expect, so a production cannot tell whether a bad shot is an audio problem or a picture one.

A named driver ends that diagnosis before it starts. This register grades the two differently for that reason alone, and the grade says nothing about which model aligns better.

2The condition is worth more than the driver

Saying performance improves with the avatar framed closer names a weak case, and the weak case is the wide two-shot that dialogue coverage is built from. That single sentence constrains a shot list, which no capability claim does.

The other entry leaves the two-speaker question untouched, so the safe assumption is alternating singles. That doubles the number of generations a conversation costs, and nothing on the page acknowledges it.

3They answer the voice question in opposite directions too

One binds a voice to an element so a character keeps it, and never says where the voice came from. The other publishes every voice with a code and an id, and never says how a voice pairs with an avatar across videos.

Persistence without selection, against selection without persistence. Between them they hold both halves of a casting system and neither publishes both.

4Each of them on its own

The column this pair was chosen for is lip-sync, and each entry has a page of its own on it. The full row for either, all five columns with the wording behind each cell, is on its model note.

  • Lip-sync
    Lip-sync is documentedstated without a mechanismKling AI, model guide / recorded 2026-09-12
  • Per-character binding
    Voices are bound to elements, so a character carries its voice between generationstied to the element systemKling AI, model guide / recorded 2026-09-12
  • Lip-sync
    Lip sync and facial expressions are generated from the spoken contentdriven by the scriptSynthesia, create an avatar / recorded 2026-09-22
  • A framing condition attached to lip-sync
    Lip sync performs best when the avatar is framed closer in the scene rather than positioned far from the cameraSynthesia, create an avatar / recorded 2026-09-22
  • Languages
    Each voice is listed with its formal and native language name, a language code, a gender, a name and a voice idpublished as a catalogue with codesSynthesia, list of supported voices / recorded 2026-09-22

5Sources

Every cell above is read from the documentation each vendor publishes, on the dates carried by the two model notes. The whole register on one column is on the field note; entries grouped by the shape of their answer are on the routes. Other pairs: SceneMixer and Synthesia, PixVerse and D-ID.