Sentioscope

Speech and voice controls, as each vendor documents them

Voice source: what a production hands over

Voice source records what a production can hand over so speech comes out as a particular character. Ten of the 17 entries answer it, between them describing a catalogue, a clone enrolled once, a clip supplied on the call, and a voice asked for in words. As of 2026-09-12.

Three answers to what a production can hand overA closed library that can be auditioned before a season starts, an open reference clip that becomes an asset to store, or nothing published at all. The field records which of the three a vendor describes.What can be handed over to decide the voiceChoose from a setA preset libraryAuditioned once,repeatable, limited towhat is in it.Supply a recordingA reference clipOpen-ended, capped inseconds, and now an assetto keep.Nothing statedNo published inputVoices exist in theoutput; their origin isnever described.Three routes, and most entries are the third
Fig. 1 A library is chosen from and a clip is supplied; the blank cell means neither is described, not that voices are absent.
Controls documented, model by modelFilled where the model documents that control, hollow where nothing is published about it.Controls documented, model by modelWhat can be suppliedWhat can besuppliedThe published wording behind the cellThe publishedwording behind th…D-IDD-ID — What can be supplied: A voice id, or a recording by urlD-ID — The published wording behind the cell: Five providers named, from Microsoft to Azure OpenAIHedraHedra — What can be supplied: A track, or a voice id from the endpointHedra — The published wording behind the cell: Audio uploaded, or speech generated inlineHeyGenHeyGen — What can be supplied: A recording to clone fromHeyGen — The published wording behind the cell: Instant from one recording, professional from twenty minutesKling AIKling AI — What can be supplied: Nothing documented as an inputKling AI — The published wording behind the cell: A voice arrives with the element, not from the userLTX StudioLTX Studio — What can be supplied: Nothing documented as an inputLTX Studio — The published wording behind the cell: Voices are attached to character ElementsLuma RayLuma Ray — What can be supplied: Nothing documented as an inputLuma Ray — The published wording behind the cell: Nothing published about voices anywhere on the pageMiniMaxMiniMax — What can be supplied: A reference clip on the callMiniMax — The published wording behind the cell: Fifteen seconds in total, across at most three clipsPixVersePixVerse — What can be supplied: A sample, or a built-in voicePixVerse — The published wording behind the cell: Built-in voices and custom voices from supplied audioRunwayRunway — What can be supplied: A 10-second to 5-minute sample, or a sentenceRunway — The published wording behind the cell: At most 10 MB, or a description of 20 charactersSceneMixerSceneMixer — What can be supplied: A preset library, or the user's own recordingsSceneMixer — The published wording behind the cell: Named in a compliance checklist rather than in product copySora 2Sora 2 — What can be supplied: Nothing documented as an inputSora 2 — The published wording behind the cell: Speakers are labelled; their voices are not specifiedsync-3sync-3 — What can be supplied: Nothing documented as an inputsync-3 — The published wording behind the cell: Text is listed as an input, voices are not discussedSynthesiaSynthesia — What can be supplied: A catalogue voice, or a cloned oneSynthesia — The published wording behind the cell: Each voice carries a language, a gender and an idVeoVeo — What can be supplied: Not documented by the vendorVeo — The published wording behind the cell: Audio is documented; its origin is not addressedViduVidu — What can be supplied: Nothing documented as an inputVidu — The published wording behind the cell: Audio is described as an output of the callWan 3.0Wan 3.0 — What can be supplied: Reference audio, 15 seconds in totalWan 3.0 — The published wording behind the cell: WAV or MP3, one to fifteen seconds a clip, 15 MBWan2.2-S2VWan2.2-S2V — What can be supplied: The audio track itselfWan2.2-S2V — The published wording behind the cell: An audio input with a reference image and a prompt
Fig. 2 Filled where the model documents that control, hollow where nothing is published about it.
Models documenting each controlHow many models document each control. A hollow column is a statement about documentation, not capability.Models documenting each controlWhat can be supplied16 of 17The published wording behind the ce…The published wording behind the cell16 of 17
Fig. 3 How many models document each control. A hollow column is a statement about documentation, not capability.
What each vendor accepts, or offers, as the origin of a voice. Recorded 2026-09-12.
ModelWhat can be suppliedThe published wording behind the cell
D-IDA voice id, or a recording by urlFive providers named, from Microsoft to Azure OpenAI
HedraA track, or a voice id from the endpointAudio uploaded, or speech generated inline
HeyGenA recording to clone fromInstant from one recording, professional from twenty minutes
Kling AINothing documented as an inputA voice arrives with the element, not from the user
LTX StudioNothing documented as an inputVoices are attached to character Elements
Luma RayNothing documented as an inputNothing published about voices anywhere on the page
MiniMaxA reference clip on the callFifteen seconds in total, across at most three clips
PixVerseA sample, or a built-in voiceBuilt-in voices and custom voices from supplied audio
RunwayA 10-second to 5-minute sample, or a sentenceAt most 10 MB, or a description of 20 characters
SceneMixerA preset library, or the user's own recordingsNamed in a compliance checklist rather than in product copy
Sora 2Nothing documented as an inputSpeakers are labelled; their voices are not specified
sync-3Nothing documented as an inputText is listed as an input, voices are not discussed
SynthesiaA catalogue voice, or a cloned oneEach voice carries a language, a gender and an id
VeoNot documented by the vendorAudio is documented; its origin is not addressed
ViduNothing documented as an inputAudio is described as an output of the call
Wan 3.0Reference audio, 15 seconds in totalWAV or MP3, one to fifteen seconds a clip, 15 MB
Wan2.2-S2VThe audio track itselfAn audio input with a reference image and a prompt

Inclusion rule. Models whose documentation says what a production may hand over to steer the voice. A model that produces voices without accepting anything from the user is recorded as documenting no input. Order. Alphabetical by model name.

1A library and a clip are not the same offer

Choosing from a preset library and supplying a recording answer the same production question in opposite directions. A library is a closed set that can be auditioned before a season starts, so casting happens once and the choice is repeatable. A reference clip is an open set, so the voice is whatever the sample produced and the sample becomes an asset to be stored.

The register keeps both in one field because a reader is asking where the voice comes from, not which mechanism the vendor prefers. What the field cannot do is rank them: a library that does not contain the voice a scene needs is worse than a clip that does.

2Fifteen seconds is the figure that shapes the casting

One entry here publishes a number rather than a capability: reference audio capped at fifteen seconds in total, split across no more than three clips. That is enough to carry a timbre and not enough to carry a reading, so the clip should be chosen for the voice rather than for the best take of the line.

The cap also sits on the call, not on the project, so a scene with several speaking parts divides the same budget. Nothing in that documentation describes storing the result, which means the clip goes in again every time the character speaks.

3The field where consent turns up

The clearest sentence in this column came out of a compliance checklist rather than a feature page: preset voices or the user's own recordings, with cloning an identifiable voice described as a likeness question and US digital-replica statutes named beside it. Obligation produced the specificity that marketing had not.

That is worth noting as a pattern for the rest of the register. Where a vendor has a reason to write something down, the writing is precise; where the only reason is to sell a capability, the wording stops at the capability and this field stays empty.

4One entry at a time on this column

A cell gets a page of its own where the vendor says something specific in it, or where its silence is unusual among the entries answering the same way. The remaining cells are left in the table above, because a page repeating one short phrase would be worse than a row carrying it.

Entries that take a clip on the call that makes it:

  • MiniMax — a reference clip on the call.
  • Wan 3.0 — reference audio, 15 seconds in total.
  • Wan2.2-S2V — the audio track itself.

Entries that build a voice once and address it afterwards by name:

  • HeyGen — a recording to clone from.
  • Runway — a 10-second to 5-minute sample, or a sentence.

Entries that offer a catalogue to choose from:

  • D-ID — a voice id, or a recording by url.
  • Hedra — a track, or a voice id from the endpoint.
  • PixVerse — a sample, or a built-in voice.
  • SceneMixer — a preset library, or the user's own recordings.
  • Synthesia — a catalogue voice, or a cloned one.

Entries that publish nothing a production could hand over:

  • Kling AI — nothing documented as an input.
  • LTX Studio — nothing documented as an input.
  • sync-3 — nothing documented as an input.
  • Voice source
    The guide tells users to use a preset voice library or their own recordingstwo routes namedSceneMixer, compliance guide / recorded 2026-09-12
  • Voice source
    Reference audio is capped at 15 seconds in total across at most 3 clipsa hard published limitMiniMax, video generation guide / recorded 2026-09-12
  • What the vendor rules out
    Cloning an identifiable voice is described as a likeness issue, with US digital-replica statutes namedSceneMixer, compliance guide / recorded 2026-09-12
  • Voice source
    A custom voice can be built from an audio sample between 10 seconds and 5 minutes long and at most 10 MBa published sample windowRunway, custom voices reference / recorded 2026-09-22
  • Voice source
    A voice can instead be asked for in words, with the description required to run at least 20 charactersa route that needs no recordingRunway, custom voices reference / recorded 2026-09-22
  • Voice source
    Reference audio is WAV or MP3, one to fifteen seconds a clip, fifteen seconds in total and no more than 15 MBa published capAlibaba Cloud, Wan 3.0 API reference / recorded 2026-09-22
  • Voice source
    A voice clone is instant from one recording or professional from twenty minutes or more, and is then passed as a voice idtwo enrolment routes with a thresholdHeyGen, API quick start / recorded 2026-09-22
  • Voice source
    Text to speech supports both built-in voices and custom voices created from user-provided sample audiocatalogue or supplied samplePixVerse, speech and lip sync guide / recorded 2026-09-22
  • Voice source
    Voice cloning carries a language list of its own, running from Afrikaans to Zulu and described as the same list used for standalone cloninga cloned voice with its own reachSynthesia, create an avatar / recorded 2026-09-22
  • Voice source
    Speech can be generated inline instead of uploaded, by naming a voice id drawn from the voices endpointa catalogue behind an endpointHedra, avatar video guide / recorded 2026-09-22

5Sources

Each cell is read from the vendor page it links to, checked 2026-09-12. The fields sit side by side on the speech table, and what counts as documented is set out on how read. The other fields: Per-character binding, Languages. All of them: the field notes.