Which answers here can be checked by software
A claim a person has to read and a claim software can check are different assets. Four entries publish something checkable: codes on voices, ids for stored voices, accepted containers and byte limits. The rest publish sentences. As of 2026-09-12.
| Model | What is documented | Detail |
|---|---|---|
| PixVerse | A speaker id, and caps in seconds and megabytes | Both sides of the request |
| Runway | A voice id, a byte ceiling and a sample window | After an asynchronous create |
| Synthesia | A language code, a gender and an id per voice | Filterable before any commitment |
| Wan 3.0 | WAV or MP3, with a fifteen-megabyte ceiling | Reference audio only |
Inclusion rule. Entries publishing at least one detail in a form software can compare: a code, an identifier, a container name or a numeric limit. A capability described in prose does not earn a row. Order. Alphabetical by model name.
1A code removes the ambiguity that ruins a language claim
A language name can mean a market, a script or a dialect depending on who wrote it. A code picks one, and it can be matched automatically against a distribution plan. That is why a catalogue with codes sits at the top of the language column even though it publishes no total at all.
It also makes an inventory auditable over time. A code that disappears from a catalogue is visible to a script; a voice quietly replaced behind a claim of broad support is visible to nobody.
2An identifier is the cheapest continuity mechanism there is
An id is short, stable and impossible to drift, so a production that records it gets a repeatable voice without storing any audio. Where a vendor publishes one, the continuity problem becomes a spreadsheet problem, which is a large improvement on a listening problem.
What none of these vendors publishes is a lifetime for the id. Whether it survives a catalogue reorganisation, a plan change or a model version is unstated everywhere, so the safe habit is to treat every id as immutable and create another rather than re-point one.
3Formats and byte ceilings prevent a whole class of wasted day
Naming the accepted containers and the maximum size removes the failures a team otherwise discovers by rejection. Only one entry publishes both for reference audio, and the numbers are small enough that a phone recording will pass.
It is a modest kind of documentation and it is the kind this register finds rarest. Vendors describe what their models can do far more readily than what their endpoints will accept.
- LanguagesEach voice is listed with its formal and native language name, a language code, a gender, a name and a voice idpublished as a catalogue with codes
- Voice sourceA custom voice can be built from an audio sample between 10 seconds and 5 minutes long and at most 10 MBa published sample window
- Per-character bindingA voice is created asynchronously, reaches a ready state with a preview, and is then referred to by its own ida stored voice object
- Voice sourceReference audio is WAV or MP3, one to fifteen seconds a clip, fifteen seconds in total and no more than 15 MBa published cap
- Per-character bindingThe chosen speaker id is passed as the lip-sync speaker on the generation requesta speaker id per call
4Sources
Every line is read from the vendor documentation linked on the speech controls page, checked 2026-09-12. What counts as documented is on how read. Nearby: What silence costs, Where a line goes, Auditioning a voice.