Sentioscope

Speech and voice controls, as each vendor documents them

PixVerse: multiple languages, and three kinds of read

The guide says multiple languages and audio types are supported, and then does something no other entry does: it names the types. Speech, singing and advertisements are three different registers, and the languages behind them are still unnamed. As of 2026-09-22.

Naming kinds of performance instead of naming marketsSpeech, singing and advertisements are three registers with different timing, breath and tolerance for drift. No other entry in this register names them, and none of the languages behind them is named here either.What the sentence namesWhat it leaves outRegistersSpeech, singing and advertisementsHow each one is priced or cappedLanguagesMultipleAny single one of themVoicesBuilt-in, or custom from a sampleWhich languages a custom voice reachesLengthSixty seconds either sideWhether a sung take fits as comfortablyA genre answer where a geography answer was asked
Fig. 1 Useful information in the wrong column: it answers whether a musical number is in scope, and not whether a series can ship in German.
PixVerse on languages, statement by statement. Read from the vendor's speech and lip sync guide on 2026-09-22.
What the documentation settlesWhat it leaves to a take
Multiple languages and audio types are supported, including speech, singing and advertisementsWhich languages, since the sentence names only the types
Text to speech supports built-in voices and custom voices from supplied sample audioWhether a custom voice carries more than one of those languages
Audio and video are each capped at sixty seconds and one hundred megabytesWhether a sung take fits inside a minute as comfortably as a spoken one

Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.

1Naming registers instead of markets

Every other entry that addresses this column is answering a geography question. This one answers a performance question: the endpoint claims to handle a sung line and read copy as well as dialogue. Those have different timing, different breath and different tolerance for drift.

It is genuinely useful information and it is in the wrong column for a market decision. A producer asking whether a series can ship in German gets no help; a producer asking whether a musical number is in scope gets an answer nobody else here offers.

2Multiple is the weakest possible total

Multiple is not a number, so it cannot even be compared with five or with ninety-five. It sits at the bottom of the grades this column keeps: a list, a count, a claim, a silence. Being a claim is better than silence only because it tells a reader the capability is intended.

Where the sentence does earn its place is beside the sixty-second ceiling. A minute of sung delivery in an unnamed language is a very specific promise, and the two figures together describe a product aimed at short pieces rather than at seasons.

3The others that claim languages without naming any

Two more entries claim breadth and name nothing. One delegates the enumeration to outside providers; the other calls the support full and leaves it there.

  • D-ID — a language field, no list.
  • Hedra — a claim with no names.
  • Languages
    Multiple languages and audio types are supported, including speech, singing, and advertisementsmultiple, none of them namedPixVerse, speech and lip sync guide / recorded 2026-09-22
  • Voice source
    Text to speech supports both built-in voices and custom voices created from user-provided sample audiocatalogue or supplied samplePixVerse, speech and lip sync guide / recorded 2026-09-22

4Sources

Read from the speech and lip sync guide at docs.platform.pixverse.ai on 2026-09-22. The same column across every entry is on languages; everything this vendor publishes about speech is on PixVerse. What counts as documented is on how read.