Whether a voice can be heard before a shot exists
Casting is normally a listening decision. Across this register exactly one entry documents something to listen to before a shot is made: a voice that reaches a ready state and arrives with a preview. As of 2026-09-12.
| Model | What is documented | Detail |
|---|---|---|
| HeyGen | A clone, enrolled and then used by id | No preview step documented |
| Runway | A ready state with a preview | Before any shot is generated |
| Synthesia | A catalogue with names, codes and ids | No sample documented on the page read |
| Wan 3.0 | Reference audio supplied on the call | The result is heard with the shot |
Inclusion rule. Entries whose documentation describes a step that returns audio for approval before a generation. A catalogue that can be browsed without a documented sample does not earn a row. Order. Alphabetical by model name.
1A preview is the difference between choosing and discovering
Where a voice can be heard first, casting is a decision somebody takes deliberately. Where it cannot, the first generation of each character is the audition, and whatever comes back is what the series has, because nothing describes a way to ask for a second opinion.
An asynchronous create is what makes the preview possible. Something has to exist between the request and the shot for there to be anything to play, and only one vendor documents that intermediate object.
2Browsing a catalogue is not the same as hearing one
A published catalogue with language codes and ids is the strongest inventory in this register and it is a list rather than a set of samples. A reader can shortlist by market and gender, and still cannot tell which of two voices suits a dramatic read.
Most platforms with catalogues almost certainly have samples somewhere in a user interface. This register records the documentation a production reads before committing, and a sample player that is not documented cannot be planned around.
3Everywhere else, the shot is the audition
Thirteen entries publish nothing on this at all. For the ones that take a reference clip, the nearest thing to an audition is generating a short shot and listening to it, which is a paid test rather than a documented step.
That is worth budgeting for explicitly. One short generation per character before a season starts is cheap, and it is the only way most of this register answers a question that casting cannot skip.
- Per-character bindingA voice is created asynchronously, reaches a ready state with a preview, and is then referred to by its own ida stored voice object
- Voice sourceA custom voice can be built from an audio sample between 10 seconds and 5 minutes long and at most 10 MBa published sample window
- Voice sourceA voice clone is instant from one recording or professional from twenty minutes or more, and is then passed as a voice idtwo enrolment routes with a threshold
- LanguagesEach voice is listed with its formal and native language name, a language code, a gender, a name and a voice idpublished as a catalogue with codes
- Voice sourceReference audio is WAV or MP3, one to fifteen seconds a clip, fifteen seconds in total and no more than 15 MBa published cap
4Sources
Every line is read from the vendor documentation linked on the speech controls page, checked 2026-09-12. What counts as documented is on how read. Nearby: Outside suppliers named, The column always filled, Speech without the bed.