PixVerse: a speech endpoint with published ceilings
PixVerse documents speech and lip sync as one endpoint. Audio can be supplied or synthesised there from text, built-in and custom voices are both accepted, and audio and video are each capped at sixty seconds and a hundred megabytes. As of 2026-09-22.
| What the documentation settles | What it leaves to a take |
|---|---|
| Text to speech supports built-in voices and custom voices from supplied sample audio | How a custom voice built from a sample compares with a built-in one |
| Audio and video are each capped at sixty seconds and one hundred megabytes | How a scene longer than a minute is assembled from calls |
| Multiple languages and audio types are supported, including speech, singing and advertisements | Which languages those are, since none of them is named |
Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.
1Two ceilings on one call, and they are not the same ceiling
Sixty seconds and a hundred megabytes are both published, and either can bind first. A minute of speech is comfortably inside the byte limit; a minute of high-bitrate video is not necessarily. A production feeding long takes into this endpoint discovers which limit applies from a rejection rather than from a calculation.
The useful habit is to treat sixty seconds as the unit of work rather than as an upper bound, and to cut at a line break instead of at a limit. Scenes assembled that way survive a later change to either ceiling.
2Types of performance, not just languages
Naming singing and advertisements alongside speech is unusual in this register, and it widens what the endpoint claims to do. Sung delivery and read copy have different timing requirements from dialogue, and a vendor that names them is describing a range of registers rather than a single read.
What the sentence does not carry is a language. Multiple is the whole answer, which leaves a market question unanswered while answering a genre question nobody else here addresses at all.
3The others that synthesise before they render
Three more entries make speech in a stage of its own. They differ on whether that stage has published limits, and on whether a voice built there can be reused.
- D-ID — from a script, or a supplied url.
- HeyGen — from a speech endpoint, then rendered.
- Synthesia — from the script, or uploaded.
- Voice sourceText to speech supports both built-in voices and custom voices created from user-provided sample audiocatalogue or supplied sample
- Length per callAudio and video are each capped at sixty seconds and one hundred megabytes
- LanguagesMultiple languages and audio types are supported, including speech, singing, and advertisementsmultiple, none of them named
4Sources
Read from the speech and lip sync guide at docs.platform.pixverse.ai on 2026-09-22. The same column across every entry is on audio source; everything this vendor publishes about speech is on PixVerse. What counts as documented is on how read.