PixVerse: a speaker id, passed on the generation request
The chosen speaker id travels on the generation request as the lip-sync speaker. It is short, stable and easy to keep, and nothing published says it belongs to a character rather than to a call. As of 2026-09-22.
| What the documentation settles | What it leaves to a take |
|---|---|
| The chosen speaker id is passed as the lip-sync speaker on the generation request | Whether a custom speaker id persists between projects |
| Text to speech supports built-in voices and custom voices from supplied sample audio | Whether a custom id is versioned when the sample changes |
| Audio and video are each capped at sixty seconds and one hundred megabytes | How a scene longer than a minute keeps one speaker across calls |
Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.
1Naming the id as the lip-sync speaker is a useful detail
Most entries that pass a voice describe it as a voice. This one describes it as the lip-sync speaker, which quietly says the same identifier decides both who is heard and whose mouth moves. In a shot with two figures that is exactly the pairing a production wants made explicit.
It is also a limit. One speaker id per request suggests one speaking part per generation, which pushes a conversation into alternating calls rather than a single two-hander, and the sixty-second ceiling points the same way.
2A stable id with no stated lifetime
An id is a much better basis for consistency than a clip, because it cannot drift. What is missing is any statement about how long it lasts, whether a custom voice can be edited, and whether editing it changes the id or the voice behind it.
For a production the cheap habit is to treat an id as immutable: never re-point one at a new sample, always create another. That keeps the record honest even where the platform makes no promise.
3The others that re-establish the voice on each call
Four more entries put the voice on the request. Two pass audio, which can drift; this one and two others pass an identifier, which cannot.
- Per-character bindingThe chosen speaker id is passed as the lip-sync speaker on the generation requesta speaker id per call
- Voice sourceText to speech supports both built-in voices and custom voices created from user-provided sample audiocatalogue or supplied sample
- Length per callAudio and video are each capped at sixty seconds and one hundred megabytes
4Sources
Read from the speech and lip sync guide at docs.platform.pixverse.ai on 2026-09-22. The same column across every entry is on per-character binding; everything this vendor publishes about speech is on PixVerse. What counts as documented is on how read.