Sentioscope

Speech and voice controls, as each vendor documents them

Which answers here can be checked by software

A claim a person has to read and a claim software can check are different assets. Four entries publish something checkable: codes on voices, ids for stored voices, accepted containers and byte limits. The rest publish sentences. As of 2026-09-12.

Prose, a number, a code: three grades of checkabilityA sentence has to be read and interpreted. A number can be compared. A code can be matched automatically against a distribution plan, which is why a catalogue with codes outranks a total with none.Most useful to a scriptA code or an idMatched automatically, and auditable over time as it changes.A numberComparable between vendors, if both measure the same quantity.A container nameSettles what a production prepares before writing any code.A sentenceWhere most of this register sits: readable, and not verifiable.Least useful
Fig. 1 A code that disappears from a catalogue is visible to a script; a voice quietly replaced behind a claim is visible to nobody.
Published details concrete enough for a script to verify or match against a plan. Recorded 2026-09-12.
ModelWhat is documentedDetail
PixVerseA speaker id, and caps in seconds and megabytesBoth sides of the request
RunwayA voice id, a byte ceiling and a sample windowAfter an asynchronous create
SynthesiaA language code, a gender and an id per voiceFilterable before any commitment
Wan 3.0WAV or MP3, with a fifteen-megabyte ceilingReference audio only

Inclusion rule. Entries publishing at least one detail in a form software can compare: a code, an identifier, a container name or a numeric limit. A capability described in prose does not earn a row. Order. Alphabetical by model name.

1A code removes the ambiguity that ruins a language claim

A language name can mean a market, a script or a dialect depending on who wrote it. A code picks one, and it can be matched automatically against a distribution plan. That is why a catalogue with codes sits at the top of the language column even though it publishes no total at all.

It also makes an inventory auditable over time. A code that disappears from a catalogue is visible to a script; a voice quietly replaced behind a claim of broad support is visible to nobody.

2An identifier is the cheapest continuity mechanism there is

An id is short, stable and impossible to drift, so a production that records it gets a repeatable voice without storing any audio. Where a vendor publishes one, the continuity problem becomes a spreadsheet problem, which is a large improvement on a listening problem.

What none of these vendors publishes is a lifetime for the id. Whether it survives a catalogue reorganisation, a plan change or a model version is unstated everywhere, so the safe habit is to treat every id as immutable and create another rather than re-point one.

3Formats and byte ceilings prevent a whole class of wasted day

Naming the accepted containers and the maximum size removes the failures a team otherwise discovers by rejection. Only one entry publishes both for reference audio, and the numbers are small enough that a phone recording will pass.

It is a modest kind of documentation and it is the kind this register finds rarest. Vendors describe what their models can do far more readily than what their endpoints will accept.

  • Languages
    Each voice is listed with its formal and native language name, a language code, a gender, a name and a voice idpublished as a catalogue with codesSynthesia, list of supported voices / recorded 2026-09-22
  • Voice source
    A custom voice can be built from an audio sample between 10 seconds and 5 minutes long and at most 10 MBa published sample windowRunway, custom voices reference / recorded 2026-09-22
  • Per-character binding
    A voice is created asynchronously, reaches a ready state with a preview, and is then referred to by its own ida stored voice objectRunway, custom voices reference / recorded 2026-09-22
  • Voice source
    Reference audio is WAV or MP3, one to fifteen seconds a clip, fifteen seconds in total and no more than 15 MBa published capAlibaba Cloud, Wan 3.0 API reference / recorded 2026-09-22
  • Per-character binding
    The chosen speaker id is passed as the lip-sync speaker on the generation requesta speaker id per callPixVerse, speech and lip sync guide / recorded 2026-09-22

4Sources

Every line is read from the vendor documentation linked on the speech controls page, checked 2026-09-12. What counts as documented is on how read. Nearby: What silence costs, Where a line goes, Auditioning a voice.