Sentioscope

Speech and voice controls, as each vendor documents them

Where the voice comes from, model by model

Eight entries generate the sound in the same pass as the picture. Seven are handed a track, or a script to read out, and asked to make a face agree with it. Two of the 17 models publish nothing about sound at all. One publishes a language catalogue with codes; the rest publish a number or nothing. As of 2026-09-22.

What each model documents about generated speechFilled = the vendor documents something in that field; hollow = nothing published. A hollow cell is a fact about the documentation, not about the model. The documented phrasing for every cell is listed under the grid.What each model documents about generated speechAudio sourceLanguagesVoice per characterVoice percharacterLip-syncD-IDD-ID — Audio source: A script or an audio urlD-ID — Languages: A language field, no listD-ID — Voice per character: A voice id from five providersD-ID — Lip-sync: Not documentedHedraHedra — Audio source: Supplied to the modelHedra — Languages: Multi-language, none namedHedra — Voice per character: Uploaded track or a voice idHedra — Lip-sync: Follows the supplied audioHeyGenHeyGen — Audio source: A separate speech endpointHeyGen — Languages: Thirty or more, countedHeyGen — Voice per character: A clone reused by idHeyGen — Lip-sync: Named alongside translationKling AIKling AI — Audio source: With the picture, VIDEO 3.0Kling AI — Languages: Five, not namedKling AI — Voice per character: Bound to elementsKling AI — Lip-sync: Documented, no mechanismLTX StudioLTX Studio — Audio source: With the picture, plus audio-to-videoLTX Studio — Languages: Not documentedLTX Studio — Voice per character: On character ElementsLTX Studio — Lip-sync: Not documentedLuma RayLuma Ray — Audio source: Not documentedLuma Ray — Languages: Not documentedLuma Ray — Voice per character: Not documentedLuma Ray — Lip-sync: Not documentedMiniMaxMiniMax — Audio source: With the pictureMiniMax — Languages: Not documentedMiniMax — Voice per character: Reference audio, 15 s per callMiniMax — Lip-sync: Tied to the on-screen speakerPixVersePixVerse — Audio source: A speech endpointPixVerse — Languages: Multiple, none namedPixVerse — Voice per character: Built-in or custom speaker idPixVerse — Lip-sync: The endpoint's whole purposeRunwayRunway — Audio source: Not documentedRunway — Languages: Not documentedRunway — Voice per character: A stored voice with an idRunway — Lip-sync: Not documentedSceneMixerSceneMixer — Audio source: With the picture, in the chosen languageSceneMixer — Languages: 15, named, plus CantoneseSceneMixer — Voice per character: Preset library or own recordingsSceneMixer — Lip-sync: No patch; the performance is in-languageSora 2Sora 2 — Audio source: With the pictureSora 2 — Languages: Not documentedSora 2 — Voice per character: Speaker labels in the promptSora 2 — Lip-sync: Long speeches said not to syncsync-3sync-3 — Audio source: Supplied or read from textsync-3 — Languages: Ninety-five or more, countedsync-3 — Voice per character: Not documentedsync-3 — Lip-sync: Matched to the audioSynthesiaSynthesia — Audio source: From the script, or uploadedSynthesia — Languages: A catalogue with codesSynthesia — Voice per character: Catalogue or a cloned voiceSynthesia — Lip-sync: From the spoken contentVeoVeo — Audio source: With the pictureVeo — Languages: Not documentedVeo — Voice per character: Not documentedVeo — Lip-sync: Not documentedViduVidu — Audio source: With the picture, speech-only availableVidu — Languages: Not documentedVidu — Voice per character: Not documentedVidu — Lip-sync: Not documentedWan 3.0Wan 3.0 — Audio source: On by defaultWan 3.0 — Languages: Not documentedWan 3.0 — Voice per character: Reference audio, 15 s per callWan 3.0 — Lip-sync: Not documentedWan2.2-S2VWan2.2-S2V — Audio source: Supplied to the modelWan2.2-S2V — Languages: Not documentedWan2.2-S2V — Voice per character: The supplied trackWan2.2-S2V — Lip-sync: Synchronised to the audio
Fig. 1 Filled = the vendor documents something in that field; hollow = nothing published. A hollow cell is a fact about the documentation, not about the model. The documented phrasing for every cell is listed under the grid.
What each model documents about generated speech. Read 2026-09-12 and 2026-09-22.
ModelAudio sourceLanguagesVoice per characterLip-syncDetail
D-IDA script or an audio urlA language field, no listA voice id from five providersNot documentedD-ID
HedraSupplied to the modelMulti-language, none namedUploaded track or a voice idFollows the supplied audioHedra
HeyGenA separate speech endpointThirty or more, countedA clone reused by idNamed alongside translationHeyGen
Kling AIWith the picture, VIDEO 3.0Five, not namedBound to elementsDocumented, no mechanismKling AI
LTX StudioWith the picture, plus audio-to-videoNot documentedOn character ElementsNot documentedLTX Studio
Luma RayNot documentedNot documentedNot documentedNot documentedLuma Ray
MiniMaxWith the pictureNot documentedReference audio, 15 s per callTied to the on-screen speakerMiniMax
PixVerseA speech endpointMultiple, none namedBuilt-in or custom speaker idThe endpoint's whole purposePixVerse
RunwayNot documentedNot documentedA stored voice with an idNot documentedRunway
SceneMixerWith the picture, in the chosen language15, named, plus CantonesePreset library or own recordingsNo patch; the performance is in-languageSceneMixer
Sora 2With the pictureNot documentedSpeaker labels in the promptLong speeches said not to syncSora 2
sync-3Supplied or read from textNinety-five or more, countedNot documentedMatched to the audiosync-3
SynthesiaFrom the script, or uploadedA catalogue with codesCatalogue or a cloned voiceFrom the spoken contentSynthesia
VeoWith the pictureNot documentedNot documentedNot documentedVeo
ViduWith the picture, speech-only availableNot documentedNot documentedNot documentedVidu
Wan 3.0On by defaultNot documentedReference audio, 15 s per callNot documentedWan 3.0
Wan2.2-S2VSupplied to the modelNot documentedThe supplied trackSynchronised to the audioWan2.2-S2V

Inclusion rule. Products whose own English documentation says something about video that carries speech, whether the sound is generated with the picture or supplied to it. A field the vendor does not document is recorded as not documented rather than inferred from generated output. Order. Alphabetical by model name.

1Three architectures sit in one table, and they fail differently

Eight products here make the sound while they make the picture, so alignment arrives for free and the words are the part that resists direction. Seven take a finished recording or a script and animate a face to it, so the words are exact and the picture has to follow. Two document neither, which leaves a buyer unable to tell from the vendor which of the two shapes is being bought.

That difference decides a schedule rather than a preference. A generated-audio route can be prompted, heard and revised inside one pass. A supplied-audio route needs casting, a recording session and a delivered track before a single frame exists, and it gives a director something to approve before any rendering is paid for.

2Keeping a voice: a stored object, an id on the call, or nothing

Four entries let a voice exist before a shot does, bound to an element, attached to a character, or created as a named object with a preview to approve first. Five re-establish it on every call, from a reference clip capped at fifteen seconds or a speaker id chosen per request. The remainder publish nothing about persistence, which across forty episodes amounts to a warning.

Both documented arrangements produce a consistent voice, and they break differently. A stored voice is wrong once if it is wrong at all. A per-call clip is an opportunity to drift on every generation, and the drift is silent: nothing in a returned file records which sample produced the voice in it.

3Languages get counted far more often than they get listed

One vendor publishes rows, each voice carrying a formal language name, a native name and a language code. One names fifteen dialogue languages and adds Cantonese for speech alone. The rest offer a number without names, a claim of multi-language support, or silence. The largest figure here, ninety-five or more, belongs to a model that works from a waveform and so has no per-language voice inventory to publish in the first place.

A count and a list are different assets. A count says the capability exists somewhere inside the product; a list says whether the market a producer was asked about is inside it, which is the question that gets asked before a season is commissioned rather than after.

4Two speakers in one frame, addressed once

The commonest shot in drama went unaddressed by every vendor in the first reading of this register. One now addresses it: a prompting guide asks for speakers to be labelled consistently and turns alternated, so each line lands on the right face. Another ties lip movement to whoever is on screen, which implies a selection is being made without describing how. A third warns that lip sync works better framed closer, which is a quiet way of saying the wide two-shot is the weak case.

Three sentences, then, on the shot a dialogue scene is built out of. Everything else in the register is silent about it, and silence here is not evidence that the case is handled.

5Per model

  • D-ID — a script or an audio url, a voice id from five providers.
  • Hedra — supplied to the model, uploaded track or a voice id.
  • HeyGen — a separate speech endpoint, a clone reused by id.
  • Kling AI — with the picture, video 3.0, bound to elements.
  • LTX Studio — with the picture, plus audio-to-video, on character elements.
  • Luma Ray — not documented, not documented.
  • MiniMax — with the picture, reference audio, 15 s per call.
  • PixVerse — a speech endpoint, built-in or custom speaker id.
  • Runway — not documented, a stored voice with an id.
  • SceneMixer — with the picture, in the chosen language, preset library or own recordings.
  • Sora 2 — with the picture, speaker labels in the prompt.
  • sync-3 — supplied or read from text, not documented.
  • Synthesia — from the script, or uploaded, catalogue or a cloned voice.
  • Veo — with the picture, not documented.
  • Vidu — with the picture, speech-only available, not documented.
  • Wan 3.0 — on by default, reference audio, 15 s per call.
  • Wan2.2-S2V — supplied to the model, the supplied track.
  • Audio source
    Described as a media generation model generating videos with synced audio, with video and audio both listed as outputgenerated with the pictureOpenAI, Sora 2 model page / recorded 2026-09-22
  • Two characters in frame
    For multi-character scenes the guide asks for speakers labelled consistently and turns alternated, so each line lands on the right characteraddressed rather than skippedOpenAI, Sora 2 prompting guide / recorded 2026-09-22
  • Voice source
    A custom voice can be built from an audio sample between 10 seconds and 5 minutes long and at most 10 MBa published sample windowRunway, custom voices reference / recorded 2026-09-22
  • What choosing silence costs
    Enabling or disabling audio does not affect pricingAlibaba Cloud, Wan 3.0 API reference / recorded 2026-09-22
  • Sound
    Plans do integrate the following audio models into the Luma generative ecosystem: ElevenLabs SFX, Music and v3named as an integration to comeLuma, information for AI assistants / recorded 2026-09-22
  • Audio source
    Avatar videos are driven by audio, and the character in the image will lip-sync and move to the audio providedhanded to the model rather than made by itHedra, avatar video guide / recorded 2026-09-22
  • Voice source
    A voice clone is instant from one recording or professional from twenty minutes or more, and is then passed as a voice idtwo enrolment routes with a thresholdHeyGen, API quick start / recorded 2026-09-22
  • Languages
    Each voice is listed with its formal and native language name, a language code, a gender, a name and a voice idpublished as a catalogue with codesSynthesia, list of supported voices / recorded 2026-09-22
  • Voice source
    Text to speech supports both built-in voices and custom voices created from user-provided sample audiocatalogue or supplied samplePixVerse, speech and lip sync guide / recorded 2026-09-22
  • Languages
    sync-3 supports 95 or more languages, described as the same coverage as the models before itpublished as a countSync, sync-3 model documentation / recorded 2026-09-22
  • Voice source
    Five speech providers are named for the voice: Microsoft, ElevenLabs, Amazon, Google and Azure OpenAIoutside catalogues named on the pageD-ID, create a talk reference / recorded 2026-09-22
  • Audio source
    The card describes audio-driven cinematic video generation from an audio input with a reference image and an optional text prompthanded to the model rather than made by itWan-AI, Wan2.2-S2V-14B model card / recorded 2026-09-22

6Sources

Each row links the page it was read from, and each entry carries its own reading date. What counts as documented is on how read; the fields themselves are described on the field notes.