Sentioscope

Speech and voice controls, as each vendor documents them

Where the voices actually come from

Four routes are documented across the register: a catalogue of ready voices, a clone enrolled from a recording, a finished track handed over, and a voice asked for in words. One vendor rules out a fifth, treating the cloning of an identifiable voice as a likeness issue. As of 2026-09-12.

The routes documented across this register, including the one ruled outTwo routes are documented: a preset library, and recordings supplied by the user. One vendor rules out a third explicitly, treating the cloning of an identifiable voice as a likeness issue and naming United States digital-replica statutes. Ruling a route out is also documentation.A preset libraryDocumentedNamed voicesSelectable again laterA supplied recordingDocumentedYour own clipShort caps, kept by youCloning a real voiceRuled outA stated refusalNamed as a likeness issueWhere a generated voice can originateA refusal is documentation too
Fig. 1 The clearest sentence about where voices come from is a refusal, and it sits in a compliance document rather than on a product page.
The routes by which a voice reaches a generated performance, as documented. Recorded 2026-09-12.
ModelWhat is documentedDetail
A finished track handed overDocumented by eightCaps run from 15 seconds to 10 minutes
A catalogue of ready voicesDocumented by five vendorsIds, sometimes with language codes
A clone from a recordingDocumented by threeOne publishes the twenty-minute threshold
A voice described in wordsDocumented by oneAt least twenty characters of description
Cloning an identifiable voiceRuled out by one vendorFramed as a likeness issue

Inclusion rule. Routes described on vendor pages. A route this register has heard about but cannot find documented is not listed. Order. By how many vendors document the route.

1The most useful sentence sits in a compliance document, not a product page

The statement that voices come from a preset library or from the user's own recordings appears in a checklist about likeness and publishing risk rather than in feature documentation. That is an odd place for a capability statement and it is the clearest one in this register.

It confirms two routes exist and stops there: it does not say how many voices the library holds, which languages they cover, or how one is attached to a character. Reading further into it than that would be inventing the product documentation the vendor has not published.

2Ruling a route out is also documentation

Most registers record what a product does. A published boundary is worth as much: knowing that a vendor treats cloning an identifiable voice as out of bounds tells a production which conversations not to start, and it is the only place in this register where anyone names a legal regime.

This site records that the boundary was stated and by whom. It does not evaluate the legal reasoning, which belongs to people qualified to do it and not to a register of published statements.

3A supplied recording has a shape the documentation rarely gives

One vendor publishes the cap, and it is tight: 15 seconds across three clips. The others that accept supplied audio say nothing about length, format or how the sample is used, so a production planning around it has one published constraint and several unknowns.

The unknowns matter for casting: whether a sample can be re-used, whether it is stored, and whether the same sample yields the same voice next month are all unanswered.

  • Voice source
    The guide tells users to use a preset voice library or their own recordingstwo routes namedSceneMixer, compliance guide / recorded 2026-09-12
  • What the vendor rules out
    Cloning an identifiable voice is described as a likeness issue, with US digital-replica statutes namedSceneMixer, compliance guide / recorded 2026-09-12
  • Voice source
    Reference audio is capped at 15 seconds in total across at most 3 clipsa hard published limitMiniMax, video generation guide / recorded 2026-09-12
  • Voice source
    A voice can instead be asked for in words, with the description required to run at least 20 charactersa route that needs no recordingRunway, custom voices reference / recorded 2026-09-22
  • Voice source
    A voice clone is instant from one recording or professional from twenty minutes or more, and is then passed as a voice idtwo enrolment routes with a thresholdHeyGen, API quick start / recorded 2026-09-22
  • Voice source
    Five speech providers are named for the voice: Microsoft, ElevenLabs, Amazon, Google and Azure OpenAIoutside catalogues named on the pageD-ID, create a talk reference / recorded 2026-09-22
  • Voice source
    Text to speech supports both built-in voices and custom voices created from user-provided sample audiocatalogue or supplied samplePixVerse, speech and lip sync guide / recorded 2026-09-22
  • Voice source
    Speech can be generated inline instead of uploaded, by naming a voice id drawn from the voices endpointa catalogue behind an endpointHedra, avatar video guide / recorded 2026-09-22
  • Voice source
    Reference audio is WAV or MP3, one to fifteen seconds a clip, fifteen seconds in total and no more than 15 MBa published capAlibaba Cloud, Wan 3.0 API reference / recorded 2026-09-22

4Sources

Every line is read from the vendor documentation linked on the speech controls page, checked 2026-09-12. What counts as documented is on how read. Nearby: Sound by default, Which language, Two in frame.