Sentioscope

Speech and voice controls, as each vendor documents them

PixVerse: a speech endpoint with published ceilings

PixVerse documents speech and lip sync as one endpoint. Audio can be supplied or synthesised there from text, built-in and custom voices are both accepted, and audio and video are each capped at sixty seconds and a hundred megabytes. As of 2026-09-22.

Two published ceilings, and either can bind firstAudio and video are each capped at sixty seconds and one hundred megabytes. A minute of speech sits well inside the byte limit; a minute of high-bitrate video need not, so the binding ceiling depends on the material.Sixty secondsOne hundred megabytesWhat it limitsHow much performance fits one callHow heavy the file may beBinds first onDialogue, which is light and longHigh-bitrate picture, which is short andheavyHow you discover itBy counting words before the callBy a rejection after itApplies toAudio and video alikeAudio and video alikePublished for both sides of the same request
Fig. 1 Treating a minute as the unit of work, and cutting at a line break rather than at a limit, survives a later change to either figure.
PixVerse on audio source, statement by statement. Read from the vendor's speech and lip sync guide on 2026-09-22.
What the documentation settlesWhat it leaves to a take
Text to speech supports built-in voices and custom voices from supplied sample audioHow a custom voice built from a sample compares with a built-in one
Audio and video are each capped at sixty seconds and one hundred megabytesHow a scene longer than a minute is assembled from calls
Multiple languages and audio types are supported, including speech, singing and advertisementsWhich languages those are, since none of them is named

Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.

1Two ceilings on one call, and they are not the same ceiling

Sixty seconds and a hundred megabytes are both published, and either can bind first. A minute of speech is comfortably inside the byte limit; a minute of high-bitrate video is not necessarily. A production feeding long takes into this endpoint discovers which limit applies from a rejection rather than from a calculation.

The useful habit is to treat sixty seconds as the unit of work rather than as an upper bound, and to cut at a line break instead of at a limit. Scenes assembled that way survive a later change to either ceiling.

2Types of performance, not just languages

Naming singing and advertisements alongside speech is unusual in this register, and it widens what the endpoint claims to do. Sung delivery and read copy have different timing requirements from dialogue, and a vendor that names them is describing a range of registers rather than a single read.

What the sentence does not carry is a language. Multiple is the whole answer, which leaves a market question unanswered while answering a genre question nobody else here addresses at all.

3The others that synthesise before they render

Three more entries make speech in a stage of its own. They differ on whether that stage has published limits, and on whether a voice built there can be reused.

  • D-ID — from a script, or a supplied url.
  • HeyGen — from a speech endpoint, then rendered.
  • Synthesia — from the script, or uploaded.
  • Voice source
    Text to speech supports both built-in voices and custom voices created from user-provided sample audiocatalogue or supplied samplePixVerse, speech and lip sync guide / recorded 2026-09-22
  • Length per call
    Audio and video are each capped at sixty seconds and one hundred megabytesPixVerse, speech and lip sync guide / recorded 2026-09-22
  • Languages
    Multiple languages and audio types are supported, including speech, singing, and advertisementsmultiple, none of them namedPixVerse, speech and lip sync guide / recorded 2026-09-22

4Sources

Read from the speech and lip sync guide at docs.platform.pixverse.ai on 2026-09-22. The same column across every entry is on audio source; everything this vendor publishes about speech is on PixVerse. What counts as documented is on how read.