PixVerse: lip sync is the endpoint's name, not its method
Lip sync is what the endpoint is called, and the chosen speaker id is passed as the lip-sync speaker. That is a routing detail rather than a mechanism: it says whose mouth, and not what the mouth follows. As of 2026-09-22.
| What the documentation settles | What it leaves to a take |
|---|---|
| The endpoint is documented for speech and lip sync together | What the movement is generated from |
| The chosen speaker id is passed as the lip-sync speaker on the request | Whether more than one speaker can be handled in one generation |
| Audio and video are each capped at sixty seconds and one hundred megabytes | Whether alignment degrades towards the end of a minute |
Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.
1Naming the speaker is a partial answer to the hard case
Most entries that name lip-sync say nothing about which face is involved. Passing a speaker id as the lip-sync speaker at least identifies the subject, which is half of the two-speaker problem. The other half, what happens to the second face, is not addressed.
One id per request suggests one speaking part per generation, so a conversation is assembled from alternating calls. That is a workable pattern and it means the shot list carries the turn-taking rather than the prompt.
2An endpoint name is the weakest kind of claim
A product can call an endpoint anything. Treating the name as documentation of a capability would put this cell on the same footing as one that describes a driver, and the two are not comparable: a name cannot be tested against a mechanism.
So the cell records the endpoint's purpose and marks the driver absent. Combined with the sixty-second ceiling and the supported audio types, it is enough to plan a short piece and not enough to predict how the face will behave in it.
3The others that name the feature and stop
Three more entries name lip-sync and leave the driver out. This is the only one where the claim is carried by the name of the endpoint itself.
- HeyGen — named, alongside translation.
- Kling AI — named, with nothing driving it.
- Sora 2 — named only where it fails.
- Voice sourceText to speech supports both built-in voices and custom voices created from user-provided sample audiocatalogue or supplied sample
- Per-character bindingThe chosen speaker id is passed as the lip-sync speaker on the generation requesta speaker id per call
- Length per callAudio and video are each capped at sixty seconds and one hundred megabytes
4Sources
Read from the speech and lip sync guide at docs.platform.pixverse.ai on 2026-09-22. The same column across every entry is on lip-sync; everything this vendor publishes about speech is on PixVerse. What counts as documented is on how read.