PixVerse: multiple languages, and three kinds of read
The guide says multiple languages and audio types are supported, and then does something no other entry does: it names the types. Speech, singing and advertisements are three different registers, and the languages behind them are still unnamed. As of 2026-09-22.
| What the documentation settles | What it leaves to a take |
|---|---|
| Multiple languages and audio types are supported, including speech, singing and advertisements | Which languages, since the sentence names only the types |
| Text to speech supports built-in voices and custom voices from supplied sample audio | Whether a custom voice carries more than one of those languages |
| Audio and video are each capped at sixty seconds and one hundred megabytes | Whether a sung take fits inside a minute as comfortably as a spoken one |
Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.
1Naming registers instead of markets
Every other entry that addresses this column is answering a geography question. This one answers a performance question: the endpoint claims to handle a sung line and read copy as well as dialogue. Those have different timing, different breath and different tolerance for drift.
It is genuinely useful information and it is in the wrong column for a market decision. A producer asking whether a series can ship in German gets no help; a producer asking whether a musical number is in scope gets an answer nobody else here offers.
2Multiple is the weakest possible total
Multiple is not a number, so it cannot even be compared with five or with ninety-five. It sits at the bottom of the grades this column keeps: a list, a count, a claim, a silence. Being a claim is better than silence only because it tells a reader the capability is intended.
Where the sentence does earn its place is beside the sixty-second ceiling. A minute of sung delivery in an unnamed language is a very specific promise, and the two figures together describe a product aimed at short pieces rather than at seasons.
3The others that claim languages without naming any
Two more entries claim breadth and name nothing. One delegates the enumeration to outside providers; the other calls the support full and leaves it there.
- LanguagesMultiple languages and audio types are supported, including speech, singing, and advertisementsmultiple, none of them named
- Voice sourceText to speech supports both built-in voices and custom voices created from user-provided sample audiocatalogue or supplied sample
4Sources
Read from the speech and lip sync guide at docs.platform.pixverse.ai on 2026-09-22. The same column across every entry is on languages; everything this vendor publishes about speech is on PixVerse. What counts as documented is on how read.