Where the voices actually come from
Four routes are documented across the register: a catalogue of ready voices, a clone enrolled from a recording, a finished track handed over, and a voice asked for in words. One vendor rules out a fifth, treating the cloning of an identifiable voice as a likeness issue. As of 2026-09-12.
| Model | What is documented | Detail |
|---|---|---|
| A finished track handed over | Documented by eight | Caps run from 15 seconds to 10 minutes |
| A catalogue of ready voices | Documented by five vendors | Ids, sometimes with language codes |
| A clone from a recording | Documented by three | One publishes the twenty-minute threshold |
| A voice described in words | Documented by one | At least twenty characters of description |
| Cloning an identifiable voice | Ruled out by one vendor | Framed as a likeness issue |
Inclusion rule. Routes described on vendor pages. A route this register has heard about but cannot find documented is not listed. Order. By how many vendors document the route.
1The most useful sentence sits in a compliance document, not a product page
The statement that voices come from a preset library or from the user's own recordings appears in a checklist about likeness and publishing risk rather than in feature documentation. That is an odd place for a capability statement and it is the clearest one in this register.
It confirms two routes exist and stops there: it does not say how many voices the library holds, which languages they cover, or how one is attached to a character. Reading further into it than that would be inventing the product documentation the vendor has not published.
2Ruling a route out is also documentation
Most registers record what a product does. A published boundary is worth as much: knowing that a vendor treats cloning an identifiable voice as out of bounds tells a production which conversations not to start, and it is the only place in this register where anyone names a legal regime.
This site records that the boundary was stated and by whom. It does not evaluate the legal reasoning, which belongs to people qualified to do it and not to a register of published statements.
3A supplied recording has a shape the documentation rarely gives
One vendor publishes the cap, and it is tight: 15 seconds across three clips. The others that accept supplied audio say nothing about length, format or how the sample is used, so a production planning around it has one published constraint and several unknowns.
The unknowns matter for casting: whether a sample can be re-used, whether it is stored, and whether the same sample yields the same voice next month are all unanswered.
- Voice sourceThe guide tells users to use a preset voice library or their own recordingstwo routes named
- What the vendor rules outCloning an identifiable voice is described as a likeness issue, with US digital-replica statutes named
- Voice sourceReference audio is capped at 15 seconds in total across at most 3 clipsa hard published limit
- Voice sourceA voice can instead be asked for in words, with the description required to run at least 20 charactersa route that needs no recording
- Voice sourceA voice clone is instant from one recording or professional from twenty minutes or more, and is then passed as a voice idtwo enrolment routes with a threshold
- Voice sourceFive speech providers are named for the voice: Microsoft, ElevenLabs, Amazon, Google and Azure OpenAIoutside catalogues named on the page
- Voice sourceText to speech supports both built-in voices and custom voices created from user-provided sample audiocatalogue or supplied sample
- Voice sourceSpeech can be generated inline instead of uploaded, by naming a voice id drawn from the voices endpointa catalogue behind an endpoint
- Voice sourceReference audio is WAV or MP3, one to fifteen seconds a clip, fifteen seconds in total and no more than 15 MBa published cap
4Sources
Every line is read from the vendor documentation linked on the speech controls page, checked 2026-09-12. What counts as documented is on how read. Nearby: Sound by default, Which language, Two in frame.