Notes on what each model documents about speech
Each note records what one vendor publishes about generated speech, in the vendor's own wording, with the page it came from and the day it was read. As of 2026-09-12.
Native audio is the first thing a note establishes. A model that produces sound along with the picture is solving a different problem from one that renders silent footage for dubbing, and the two are described with similar enthusiasm in launch material.
Reference audio is the second. Supplying a voice sample is how a character keeps the same voice across a season, and the documentation for it ranges from a detailed parameter list to a single sentence to nothing at all.
Timing is the field vendors address least. A line that runs longer than the shot it belongs to is an editing problem, and nothing in most documentation says whether duration can be controlled, hinted at, or only discovered after the render.
1The notes
- D-ID — What D-ID documents about generated speech: where the audio comes from, how a voice attaches to a character, and whether lip-sync is covered.
- Hedra — What Hedra documents about generated speech: where the audio comes from, how a voice attaches to a character, and whether lip-sync is covered.
- HeyGen — What HeyGen documents about generated speech: where the audio comes from, how a voice attaches to a character, and whether lip-sync is covered.
- Kling AI — What Kling AI documents about generated speech: where the audio comes from, how a voice attaches to a character, and whether lip-sync is covered.
- LTX Studio — What LTX Studio documents about generated speech: where the audio comes from, how a voice attaches to a character, and whether lip-sync is covered.
- Luma Ray — What Luma Ray documents about generated speech: where the audio comes from, how a voice attaches to a character, and whether lip-sync is covered.
- MiniMax — What MiniMax documents about generated speech: where the audio comes from, how a voice attaches to a character, and whether lip-sync is covered.
- PixVerse — What PixVerse documents about generated speech: where the audio comes from, how a voice attaches to a character, and whether lip-sync is covered.
- Runway — What Runway documents about generated speech: where the audio comes from, how a voice attaches to a character, and whether lip-sync is covered.
- SceneMixer — What SceneMixer documents about generated speech: where the audio comes from, how a voice attaches to a character, and whether lip-sync is covered.
- Sora 2 — What Sora 2 documents about generated speech: where the audio comes from, how a voice attaches to a character, and whether lip-sync is covered.
- sync-3 — What sync-3 documents about generated speech: where the audio comes from, how a voice attaches to a character, and whether lip-sync is covered.
- Synthesia — What Synthesia documents about generated speech: where the audio comes from, how a voice attaches to a character, and whether lip-sync is covered.
- Veo — What Veo documents about generated speech: where the audio comes from, how a voice attaches to a character, and whether lip-sync is covered.
- Vidu — What Vidu documents about generated speech: where the audio comes from, how a voice attaches to a character, and whether lip-sync is covered.
- Wan 3.0 — What Wan 3.0 documents about generated speech: where the audio comes from, how a voice attaches to a character, and whether lip-sync is covered.
- Wan2.2-S2V — What Wan2.2-S2V documents about generated speech: where the audio comes from, how a voice attaches to a character, and whether lip-sync is covered.
A model earns a note when its own documentation addresses speech at all, including when it addresses it to say audio is not produced.
Other notes: Speech controls, Fields, Routes, Side by side, Questions, Terms, Learn, Data. What counts as a documented control is set out on the reading page.