Sentioscope

Speech and voice controls, as each vendor documents them

PixVerse: lip sync is the endpoint's name, not its method

Lip sync is what the endpoint is called, and the chosen speaker id is passed as the lip-sync speaker. That is a routing detail rather than a mechanism: it says whose mouth, and not what the mouth follows. As of 2026-09-22.

A speaker named, and a mechanism still missingPassing a speaker id as the lip-sync speaker identifies the subject, which is half of the two-speaker problem. What happens to the second face, and what the movement is generated from, are both unaddressed.AnsweredNot answeredWhose mouthThe speaker id on the requestWhat the second face doesHow many at onceOne id, so one speaking partHow a conversation is assembledWhat drives itNothing statedAudio, text, or a later passHow longSixty seconds either sideWhether it degrades towards the endAn endpoint name is the weakest kind of claim
Fig. 1 A product can call an endpoint anything, so treating the name as documentation would put this cell on a footing it has not earned.
PixVerse on lip-sync, statement by statement. Read from the vendor's speech and lip sync guide on 2026-09-22.
What the documentation settlesWhat it leaves to a take
The endpoint is documented for speech and lip sync togetherWhat the movement is generated from
The chosen speaker id is passed as the lip-sync speaker on the requestWhether more than one speaker can be handled in one generation
Audio and video are each capped at sixty seconds and one hundred megabytesWhether alignment degrades towards the end of a minute

Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.

1Naming the speaker is a partial answer to the hard case

Most entries that name lip-sync say nothing about which face is involved. Passing a speaker id as the lip-sync speaker at least identifies the subject, which is half of the two-speaker problem. The other half, what happens to the second face, is not addressed.

One id per request suggests one speaking part per generation, so a conversation is assembled from alternating calls. That is a workable pattern and it means the shot list carries the turn-taking rather than the prompt.

2An endpoint name is the weakest kind of claim

A product can call an endpoint anything. Treating the name as documentation of a capability would put this cell on the same footing as one that describes a driver, and the two are not comparable: a name cannot be tested against a mechanism.

So the cell records the endpoint's purpose and marks the driver absent. Combined with the sixty-second ceiling and the supported audio types, it is enough to plan a short piece and not enough to predict how the face will behave in it.

3The others that name the feature and stop

Three more entries name lip-sync and leave the driver out. This is the only one where the claim is carried by the name of the endpoint itself.

  • HeyGen — named, alongside translation.
  • Kling AI — named, with nothing driving it.
  • Sora 2 — named only where it fails.
  • Voice source
    Text to speech supports both built-in voices and custom voices created from user-provided sample audiocatalogue or supplied samplePixVerse, speech and lip sync guide / recorded 2026-09-22
  • Per-character binding
    The chosen speaker id is passed as the lip-sync speaker on the generation requesta speaker id per callPixVerse, speech and lip sync guide / recorded 2026-09-22
  • Length per call
    Audio and video are each capped at sixty seconds and one hundred megabytesPixVerse, speech and lip sync guide / recorded 2026-09-22

4Sources

Read from the speech and lip sync guide at docs.platform.pixverse.ai on 2026-09-22. The same column across every entry is on lip-sync; everything this vendor publishes about speech is on PixVerse. What counts as documented is on how read.