Sentioscope

Speech and voice controls, as each vendor documents them

What video models document about speech

Sentioscope records what each AI video model publishes about generated speech: whether the audio comes out of the model or has to be handed to it, how a voice attaches to a character, and what drives lip-sync. Seventeen models, five fields, vendor documentation only. As of 2026-09-22.

Controls documented, model by modelFilled where the model documents that control, hollow where nothing is published about it.Controls documented, model by modelAudio sourceLanguagesD-IDD-ID — Audio source: A script or an audio urlD-ID — Languages: A language field, no listHedraHedra — Audio source: Supplied to the modelHedra — Languages: Multi-language, none namedHeyGenHeyGen — Audio source: A separate speech endpointHeyGen — Languages: Thirty or more, countedKling AIKling AI — Audio source: With the picture, VIDEO 3.0Kling AI — Languages: Five, not namedLTX StudioLTX Studio — Audio source: With the picture, plus audio-to-videoLTX Studio — Languages: Not documentedLuma RayLuma Ray — Audio source: Not documentedLuma Ray — Languages: Not documentedMiniMaxMiniMax — Audio source: With the pictureMiniMax — Languages: Not documentedPixVersePixVerse — Audio source: A speech endpointPixVerse — Languages: Multiple, none namedRunwayRunway — Audio source: Not documentedRunway — Languages: Not documentedSceneMixerSceneMixer — Audio source: With the picture, in the chosen languageSceneMixer — Languages: 15, named, plus CantoneseSora 2Sora 2 — Audio source: With the pictureSora 2 — Languages: Not documentedsync-3sync-3 — Audio source: Supplied or read from textsync-3 — Languages: Ninety-five or more, countedSynthesiaSynthesia — Audio source: From the script, or uploadedSynthesia — Languages: A catalogue with codesVeoVeo — Audio source: With the pictureVeo — Languages: Not documentedViduVidu — Audio source: With the picture, speech-only availableVidu — Languages: Not documentedWan 3.0Wan 3.0 — Audio source: On by defaultWan 3.0 — Languages: Not documentedWan2.2-S2VWan2.2-S2V — Audio source: Supplied to the modelWan2.2-S2V — Languages: Not documented
Fig. 1 Filled where the model documents that control, hollow where nothing is published about it.
Models documenting each controlHow many models document each control. A hollow column is a statement about documentation, not capability.Models documenting each controlAudio source15 of 17Languages8 of 17
Fig. 2 How many models document each control. A hollow column is a statement about documentation, not capability.
Every model tracked here, and what it documents about speech. Read 2026-09-12 and 2026-09-22.
ModelAudio sourceLanguages
D-IDA script or an audio urlA language field, no list
HedraSupplied to the modelMulti-language, none named
HeyGenA separate speech endpointThirty or more, counted
Kling AIWith the picture, VIDEO 3.0Five, not named
LTX StudioWith the picture, plus audio-to-videoNot documented
Luma RayNot documentedNot documented
MiniMaxWith the pictureNot documented
PixVerseA speech endpointMultiple, none named
RunwayNot documentedNot documented
SceneMixerWith the picture, in the chosen language15, named, plus Cantonese
Sora 2With the pictureNot documented
sync-3Supplied or read from textNinety-five or more, counted
SynthesiaFrom the script, or uploadedA catalogue with codes
VeoWith the pictureNot documented
ViduWith the picture, speech-only availableNot documented
Wan 3.0On by defaultNot documented
Wan2.2-S2VSupplied to the modelNot documented

Inclusion rule. Products whose own documentation says something about video that carries speech. Nothing published about sound means no row. Order. Alphabetical by model.

1Words worth pinning down

Speech terminology is used loosely almost everywhere, so the words are pinned down here one page at a time: native audio, reference audio, voice id, language code, lip-sync and the rest of the vocabulary. Each sets out what the vendors document, with the page it came from and the day it was checked.

2What is recorded here

Sound breaks a serial in ways no single frame will reveal. A face can be kept consistent with a reference image; a voice has to be kept consistent by whatever mechanism the vendor provides, and those mechanisms differ more than the picture side does. This register writes them out rather than reducing them to a column of ticks.

Everything comes from vendor documentation. Nothing is filled in by generating a clip and listening to it, because one clip is one sample and a register built from samples records the sampling rather than the product.

3Start here

  • Speech controls — every entry, four columns, one table.
  • Sora 2 — where a line of dialogue goes in the prompt.
  • Wan 3.0 — sound on by default, and silence costs the same.
  • Hedra — audio is the input, so the take sets the length.
  • Synthesia — a language catalogue with codes in it.
  • Runway — a voice built from a sample or from a sentence.
  • sync-3 — ninety-five languages, counted not named.
  • Luma Ray — an entry where every field is empty.
  • SceneMixer — a preset library, stated in a compliance guide.
  • All model notes — the rest of the register, one page each.
  • How read — why nothing is filled in from output.

4Questions about the voice

5Explainers

How generated speech is produced and what to decide before a dialogue-heavy series starts. These pages describe mechanisms and habits; what each model documents about them sits in the speech controls table.

  • Writing the lines — how dialogue has to be written so a generated read lands where it should.
  • Dub or regenerate — two routes into another language, and what each one costs you.
  • How lip-sync works — three architectures, and why the two-speaker shot breaks all of them differently.
  • Casting a voice — picking and holding a voice, and the fifteen-second problem.
  • Three routes to a voice — a preset catalogue, a cloned voice, or speech made with the picture.

The five fields were fixed before any vendor was read, which is what keeps the table usable when one vendor publishes a great deal about one field and nothing about the rest. A model with a detailed voice system and no word about lip-sync ends up with two filled fields and two marked not documented, and the shape of that row carries information of its own: it says the vendor treats voice as a casting feature and lip-sync as an implementation detail nobody needs to plan around.

6What a not-documented cell means

It means the vendor does not publish it. Several of these products plainly generate speech in more than one language and handle more than one speaker; almost none describes how. Recording the absence is the point, because the questions a drama production asks first are the ones the documentation answers least.

Sentioscope records published vendor documentation. It does not test models and it does not rank them.

7The parts of these notes

These sections record what each model publishes about speech: the controls it names, the languages it lists, and the questions only listening could settle.

  • Data — Recorded speech and voice controls as a comma-separated file, each row carrying the…
  • Speech controls — The comparison table for documented speech controls…
  • Models — A note per model…
  • Fields — One page per field in the speech table…
  • Routes — Entries collected by how they answer one column…
  • Side by side — Pairs of entries put on one page because at least one column lands them at opposite…
  • Questions — Questions about dialogue in AI video, answered from vendor documentation rather than…
  • Terms — Definitions for the speech vocabulary these notes use, from native audio and reference…
  • Learn — Background notes on dialogue in generative video…