Sentioscope

Speech and voice controls, as each vendor documents them

Entries that name lip-sync and leave the driver out

Four entries name lip-sync and stop. One puts it in a translation sentence, one in the name of an endpoint, one as a capability of the model, and one only in a warning about where it stops holding. As of 2026-09-22.

Three mechanisms, and a claim that predicts none of themSound and picture made together tends to produce soft movement. Animating to a track produces tight movement on a face that may not match. Repairing afterwards produces a mouth that can look pasted on.What a driver statement would settleMade togetherMovement can be approximate, because nothing external sets thetiming.Animated to a trackTight timing, on a face that may not sit in the shot.Repaired afterwardsA mouth that can read as pasted onto footage it did not comefrom.Named, not explainedWhere these four sit: no way to tell which artefact to expect.And what a bare capability claim leaves open
Fig. 1 A well-aligned model with a thin page lands in this group, so the grade is about the sentence and not about the product.
Four lip-sync claims without a driver, beside what each vendor says about a voice lasting. Recorded 2026-09-22.
ModelHow the claim is madeAnd what holds a voice to a character
HeyGenNamed, alongside translationOn a stored clone
Kling AINamed, with nothing driving itOn the element
PixVerseNamed as the endpoint's purposeOn the request
Sora 2Named only where it failsOn the prompt

Inclusion rule. Entries whose documentation names lip-sync, or names where it fails, without naming what the mouth follows. Entries with a named driver are on a separate route page. Order. Alphabetical by model name.

1Three mechanisms, three visibly different artefacts

Sound and picture made together tends to produce soft or approximate movement. Animating to a supplied track produces tight movement on a face that may not match the rest of the shot. Repairing afterwards produces a mouth that can look pasted on. A claim that names none of these predicts none of them.

That is why the register grades a driver above a feature. The grade is about how much a reader can plan from a sentence, not about how good the product is, and a well-aligned model with a thin page lands in this group.

2One of these names only the failure, and that is more useful

A warning that long, complex speeches are unlikely to sync tells a writer to break a speech into exchanges. It is not a mechanism, so the field stays unstated, and it is the most actionable sentence in the group.

Being told where something stops working is often worth more than being told that it works. It converts directly into a rule for a shot list, which a capability claim never does.

3An endpoint name is the weakest claim of the four

A product can call an endpoint whatever it likes. Treating the name as documentation would put it on the same footing as a described mechanism, and the two cannot be compared because a name cannot be tested.

That entry does pass a speaker id described as the lip-sync speaker, which at least identifies whose mouth is involved. Half of the two-speaker problem answered by a routing detail is still more than the other three manage.

4The entries on this route, one page each

Each of these links to that entry on the field this route groups by. The wording behind every cell, and the date it was read, sits on the page it links to.

  • HeyGen — named, alongside translation.
  • Kling AI — named, with nothing driving it.
  • PixVerse — named as the endpoint's purpose.
  • Sora 2 — named only where it fails.
  • Lip-sync
    Lip-sync is documentedstated without a mechanismKling AI, model guide / recorded 2026-09-12
  • Languages
    Video translation is documented for thirty or more languages, with voice cloning and lip-synccounted on the translation endpointHeyGen, API quick start / recorded 2026-09-22
  • Per-character binding
    The chosen speaker id is passed as the lip-sync speaker on the generation requesta speaker id per callPixVerse, speech and lip sync guide / recorded 2026-09-22
  • Dialogue timing against clip length
    A four-second shot is said to accommodate one or two short exchanges and an eight-second clip a few more, with long speeches unlikely to syncOpenAI, Sora 2 prompting guide / recorded 2026-09-22
  • Per-character binding
    Voices are bound to elements, so a character carries its voice between generationstied to the element systemKling AI, model guide / recorded 2026-09-12

5Sources

Membership of this route is decided by the wording in the field it groups by, read from the documentation each vendor publishes. The column itself is on the field note, all five columns are on the speech table, and what counts as documented is on how read. Other routes: The subject skipped, Sound in the same pass.