Sentioscope

Speech and voice controls, as each vendor documents them

Notes on what each model documents about speech

Each note records what one vendor publishes about generated speech, in the vendor's own wording, with the page it came from and the day it was read. As of 2026-09-12.

Three ways a model can arrive at spoken dialogueSpeech can be generated together with the picture, supplied separately and matched to it, or added afterwards by a second service. The three produce different failures: drifting lip movement, timing that will not fit, and a voice that changes between episodes.Where the audio comes fromWith the pictureGenerated togetherOne pass. Lip movementtends to hold; the wordsare harder to control.Supplied to itMatched to a trackControl over the words.The picture has to follow,and sometimes does not.Added laterA second serviceFull control and aseparate voice identity tokeep stable across aseason.Three routes, three kinds of failure
Fig. 1 A vendor that names only the feature has not said which of the three it is doing, and the failure mode follows from that choice.

Native audio is the first thing a note establishes. A model that produces sound along with the picture is solving a different problem from one that renders silent footage for dubbing, and the two are described with similar enthusiasm in launch material.

Reference audio is the second. Supplying a voice sample is how a character keeps the same voice across a season, and the documentation for it ranges from a detailed parameter list to a single sentence to nothing at all.

Timing is the field vendors address least. A line that runs longer than the shot it belongs to is an editing problem, and nothing in most documentation says whether duration can be controlled, hinted at, or only discovered after the render.

1The notes

  • D-ID — What D-ID documents about generated speech: where the audio comes from, how a voice attaches to a character, and whether lip-sync is covered.
  • Hedra — What Hedra documents about generated speech: where the audio comes from, how a voice attaches to a character, and whether lip-sync is covered.
  • HeyGen — What HeyGen documents about generated speech: where the audio comes from, how a voice attaches to a character, and whether lip-sync is covered.
  • Kling AI — What Kling AI documents about generated speech: where the audio comes from, how a voice attaches to a character, and whether lip-sync is covered.
  • LTX Studio — What LTX Studio documents about generated speech: where the audio comes from, how a voice attaches to a character, and whether lip-sync is covered.
  • Luma Ray — What Luma Ray documents about generated speech: where the audio comes from, how a voice attaches to a character, and whether lip-sync is covered.
  • MiniMax — What MiniMax documents about generated speech: where the audio comes from, how a voice attaches to a character, and whether lip-sync is covered.
  • PixVerse — What PixVerse documents about generated speech: where the audio comes from, how a voice attaches to a character, and whether lip-sync is covered.
  • Runway — What Runway documents about generated speech: where the audio comes from, how a voice attaches to a character, and whether lip-sync is covered.
  • SceneMixer — What SceneMixer documents about generated speech: where the audio comes from, how a voice attaches to a character, and whether lip-sync is covered.
  • Sora 2 — What Sora 2 documents about generated speech: where the audio comes from, how a voice attaches to a character, and whether lip-sync is covered.
  • sync-3 — What sync-3 documents about generated speech: where the audio comes from, how a voice attaches to a character, and whether lip-sync is covered.
  • Synthesia — What Synthesia documents about generated speech: where the audio comes from, how a voice attaches to a character, and whether lip-sync is covered.
  • Veo — What Veo documents about generated speech: where the audio comes from, how a voice attaches to a character, and whether lip-sync is covered.
  • Vidu — What Vidu documents about generated speech: where the audio comes from, how a voice attaches to a character, and whether lip-sync is covered.
  • Wan 3.0 — What Wan 3.0 documents about generated speech: where the audio comes from, how a voice attaches to a character, and whether lip-sync is covered.
  • Wan2.2-S2V — What Wan2.2-S2V documents about generated speech: where the audio comes from, how a voice attaches to a character, and whether lip-sync is covered.

A model earns a note when its own documentation addresses speech at all, including when it addresses it to say audio is not produced.

Other notes: Speech controls, Fields, Routes, Side by side, Questions, Terms, Learn, Data. What counts as a documented control is set out on the reading page.