What video models document about speech
Sentioscope records what each AI video model publishes about generated speech: whether the audio comes out of the model or has to be handed to it, how a voice attaches to a character, and what drives lip-sync. Seventeen models, five fields, vendor documentation only. As of 2026-09-22.
| Model | Audio source | Languages |
|---|---|---|
| D-ID | A script or an audio url | A language field, no list |
| Hedra | Supplied to the model | Multi-language, none named |
| HeyGen | A separate speech endpoint | Thirty or more, counted |
| Kling AI | With the picture, VIDEO 3.0 | Five, not named |
| LTX Studio | With the picture, plus audio-to-video | Not documented |
| Luma Ray | Not documented | Not documented |
| MiniMax | With the picture | Not documented |
| PixVerse | A speech endpoint | Multiple, none named |
| Runway | Not documented | Not documented |
| SceneMixer | With the picture, in the chosen language | 15, named, plus Cantonese |
| Sora 2 | With the picture | Not documented |
| sync-3 | Supplied or read from text | Ninety-five or more, counted |
| Synthesia | From the script, or uploaded | A catalogue with codes |
| Veo | With the picture | Not documented |
| Vidu | With the picture, speech-only available | Not documented |
| Wan 3.0 | On by default | Not documented |
| Wan2.2-S2V | Supplied to the model | Not documented |
Inclusion rule. Products whose own documentation says something about video that carries speech. Nothing published about sound means no row. Order. Alphabetical by model.
1Words worth pinning down
Speech terminology is used loosely almost everywhere, so the words are pinned down here one page at a time: native audio, reference audio, voice id, language code, lip-sync and the rest of the vocabulary. Each sets out what the vendors document, with the page it came from and the day it was checked.
2What is recorded here
Sound breaks a serial in ways no single frame will reveal. A face can be kept consistent with a reference image; a voice has to be kept consistent by whatever mechanism the vendor provides, and those mechanisms differ more than the picture side does. This register writes them out rather than reducing them to a column of ticks.
Everything comes from vendor documentation. Nothing is filled in by generating a clip and listening to it, because one clip is one sample and a register built from samples records the sampling rather than the product.
3Start here
- Speech controls — every entry, four columns, one table.
- Sora 2 — where a line of dialogue goes in the prompt.
- Wan 3.0 — sound on by default, and silence costs the same.
- Hedra — audio is the input, so the take sets the length.
- Synthesia — a language catalogue with codes in it.
- Runway — a voice built from a sample or from a sentence.
- sync-3 — ninety-five languages, counted not named.
- Luma Ray — an entry where every field is empty.
- SceneMixer — a preset library, stated in a compliance guide.
- All model notes — the rest of the register, one page each.
- How read — why nothing is filled in from output.
4Questions about the voice
5Explainers
How generated speech is produced and what to decide before a dialogue-heavy series starts. These pages describe mechanisms and habits; what each model documents about them sits in the speech controls table.
- Writing the lines — how dialogue has to be written so a generated read lands where it should.
- Dub or regenerate — two routes into another language, and what each one costs you.
- How lip-sync works — three architectures, and why the two-speaker shot breaks all of them differently.
- Casting a voice — picking and holding a voice, and the fifteen-second problem.
- Three routes to a voice — a preset catalogue, a cloned voice, or speech made with the picture.
- How a character keeps the same voice — bound, re-supplied, or unstated.
- Lip-sync is answered three ways — a mechanism, a capability, a removed step.
- Where the voices come from — a library, a recording, and one route ruled out.
- Which models return sound by default — and which separates speech from the bed.
- Which language the performance is in — a named list against a count.
- Two speakers in frame — the commonest shot, addressed by one guide.
- Why API docs answer the wrong questions — parameters versus behaviour.
The five fields were fixed before any vendor was read, which is what keeps the table usable when one vendor publishes a great deal about one field and nothing about the rest. A model with a detailed voice system and no word about lip-sync ends up with two filled fields and two marked not documented, and the shape of that row carries information of its own: it says the vendor treats voice as a casting feature and lip-sync as an implementation detail nobody needs to plan around.
6What a not-documented cell means
It means the vendor does not publish it. Several of these products plainly generate speech in more than one language and handle more than one speaker; almost none describes how. Recording the absence is the point, because the questions a drama production asks first are the ones the documentation answers least.
Sentioscope records published vendor documentation. It does not test models and it does not rank them.
7The parts of these notes
These sections record what each model publishes about speech: the controls it names, the languages it lists, and the questions only listening could settle.
- Data — Recorded speech and voice controls as a comma-separated file, each row carrying the…
- Speech controls — The comparison table for documented speech controls…
- Models — A note per model…
- Fields — One page per field in the speech table…
- Routes — Entries collected by how they answer one column…
- Side by side — Pairs of entries put on one page because at least one column lands them at opposite…
- Questions — Questions about dialogue in AI video, answered from vendor documentation rather than…
- Terms — Definitions for the speech vocabulary these notes use, from native audio and reference…
- Learn — Background notes on dialogue in generative video…