Where the voice comes from, model by model
Eight entries generate the sound in the same pass as the picture. Seven are handed a track, or a script to read out, and asked to make a face agree with it. Two of the 17 models publish nothing about sound at all. One publishes a language catalogue with codes; the rest publish a number or nothing. As of 2026-09-22.
| Model | Audio source | Languages | Voice per character | Lip-sync | Detail |
|---|---|---|---|---|---|
| D-ID | A script or an audio url | A language field, no list | A voice id from five providers | Not documented | D-ID |
| Hedra | Supplied to the model | Multi-language, none named | Uploaded track or a voice id | Follows the supplied audio | Hedra |
| HeyGen | A separate speech endpoint | Thirty or more, counted | A clone reused by id | Named alongside translation | HeyGen |
| Kling AI | With the picture, VIDEO 3.0 | Five, not named | Bound to elements | Documented, no mechanism | Kling AI |
| LTX Studio | With the picture, plus audio-to-video | Not documented | On character Elements | Not documented | LTX Studio |
| Luma Ray | Not documented | Not documented | Not documented | Not documented | Luma Ray |
| MiniMax | With the picture | Not documented | Reference audio, 15 s per call | Tied to the on-screen speaker | MiniMax |
| PixVerse | A speech endpoint | Multiple, none named | Built-in or custom speaker id | The endpoint's whole purpose | PixVerse |
| Runway | Not documented | Not documented | A stored voice with an id | Not documented | Runway |
| SceneMixer | With the picture, in the chosen language | 15, named, plus Cantonese | Preset library or own recordings | No patch; the performance is in-language | SceneMixer |
| Sora 2 | With the picture | Not documented | Speaker labels in the prompt | Long speeches said not to sync | Sora 2 |
| sync-3 | Supplied or read from text | Ninety-five or more, counted | Not documented | Matched to the audio | sync-3 |
| Synthesia | From the script, or uploaded | A catalogue with codes | Catalogue or a cloned voice | From the spoken content | Synthesia |
| Veo | With the picture | Not documented | Not documented | Not documented | Veo |
| Vidu | With the picture, speech-only available | Not documented | Not documented | Not documented | Vidu |
| Wan 3.0 | On by default | Not documented | Reference audio, 15 s per call | Not documented | Wan 3.0 |
| Wan2.2-S2V | Supplied to the model | Not documented | The supplied track | Synchronised to the audio | Wan2.2-S2V |
Inclusion rule. Products whose own English documentation says something about video that carries speech, whether the sound is generated with the picture or supplied to it. A field the vendor does not document is recorded as not documented rather than inferred from generated output. Order. Alphabetical by model name.
1Three architectures sit in one table, and they fail differently
Eight products here make the sound while they make the picture, so alignment arrives for free and the words are the part that resists direction. Seven take a finished recording or a script and animate a face to it, so the words are exact and the picture has to follow. Two document neither, which leaves a buyer unable to tell from the vendor which of the two shapes is being bought.
That difference decides a schedule rather than a preference. A generated-audio route can be prompted, heard and revised inside one pass. A supplied-audio route needs casting, a recording session and a delivered track before a single frame exists, and it gives a director something to approve before any rendering is paid for.
2Keeping a voice: a stored object, an id on the call, or nothing
Four entries let a voice exist before a shot does, bound to an element, attached to a character, or created as a named object with a preview to approve first. Five re-establish it on every call, from a reference clip capped at fifteen seconds or a speaker id chosen per request. The remainder publish nothing about persistence, which across forty episodes amounts to a warning.
Both documented arrangements produce a consistent voice, and they break differently. A stored voice is wrong once if it is wrong at all. A per-call clip is an opportunity to drift on every generation, and the drift is silent: nothing in a returned file records which sample produced the voice in it.
3Languages get counted far more often than they get listed
One vendor publishes rows, each voice carrying a formal language name, a native name and a language code. One names fifteen dialogue languages and adds Cantonese for speech alone. The rest offer a number without names, a claim of multi-language support, or silence. The largest figure here, ninety-five or more, belongs to a model that works from a waveform and so has no per-language voice inventory to publish in the first place.
A count and a list are different assets. A count says the capability exists somewhere inside the product; a list says whether the market a producer was asked about is inside it, which is the question that gets asked before a season is commissioned rather than after.
4Two speakers in one frame, addressed once
The commonest shot in drama went unaddressed by every vendor in the first reading of this register. One now addresses it: a prompting guide asks for speakers to be labelled consistently and turns alternated, so each line lands on the right face. Another ties lip movement to whoever is on screen, which implies a selection is being made without describing how. A third warns that lip sync works better framed closer, which is a quiet way of saying the wide two-shot is the weak case.
Three sentences, then, on the shot a dialogue scene is built out of. Everything else in the register is silent about it, and silence here is not evidence that the case is handled.
5Per model
- D-ID — a script or an audio url, a voice id from five providers.
- Hedra — supplied to the model, uploaded track or a voice id.
- HeyGen — a separate speech endpoint, a clone reused by id.
- Kling AI — with the picture, video 3.0, bound to elements.
- LTX Studio — with the picture, plus audio-to-video, on character elements.
- Luma Ray — not documented, not documented.
- MiniMax — with the picture, reference audio, 15 s per call.
- PixVerse — a speech endpoint, built-in or custom speaker id.
- Runway — not documented, a stored voice with an id.
- SceneMixer — with the picture, in the chosen language, preset library or own recordings.
- Sora 2 — with the picture, speaker labels in the prompt.
- sync-3 — supplied or read from text, not documented.
- Synthesia — from the script, or uploaded, catalogue or a cloned voice.
- Veo — with the picture, not documented.
- Vidu — with the picture, speech-only available, not documented.
- Wan 3.0 — on by default, reference audio, 15 s per call.
- Wan2.2-S2V — supplied to the model, the supplied track.
- Audio sourceDescribed as a media generation model generating videos with synced audio, with video and audio both listed as outputgenerated with the picture
- Two characters in frameFor multi-character scenes the guide asks for speakers labelled consistently and turns alternated, so each line lands on the right characteraddressed rather than skipped
- Voice sourceA custom voice can be built from an audio sample between 10 seconds and 5 minutes long and at most 10 MBa published sample window
- What choosing silence costsEnabling or disabling audio does not affect pricing
- SoundPlans do integrate the following audio models into the Luma generative ecosystem: ElevenLabs SFX, Music and v3named as an integration to come
- Audio sourceAvatar videos are driven by audio, and the character in the image will lip-sync and move to the audio providedhanded to the model rather than made by it
- Voice sourceA voice clone is instant from one recording or professional from twenty minutes or more, and is then passed as a voice idtwo enrolment routes with a threshold
- LanguagesEach voice is listed with its formal and native language name, a language code, a gender, a name and a voice idpublished as a catalogue with codes
- Voice sourceText to speech supports both built-in voices and custom voices created from user-provided sample audiocatalogue or supplied sample
- Languagessync-3 supports 95 or more languages, described as the same coverage as the models before itpublished as a count
- Voice sourceFive speech providers are named for the voice: Microsoft, ElevenLabs, Amazon, Google and Azure OpenAIoutside catalogues named on the page
- Audio sourceThe card describes audio-driven cinematic video generation from an audio input with a reference image and an optional text prompthanded to the model rather than made by it
6Sources
Each row links the page it was read from, and each entry carries its own reading date. What counts as documented is on how read; the fields themselves are described on the field notes.