The speech table, and what each column asks of a vendor
One table holds the four fields these notes keep for every model. The table is the fastest way in; the model notes behind it carry the wording each cell was read from. As of 2026-09-12.
Four fields, chosen because they are the four a production has to settle before writing dialogue. Whether audio is generated with the picture or attached afterwards. Whether a reference voice can be supplied. Which languages the vendor names for speech. What, if anything, is documented about line timing.
The fields were not chosen to flatter any model. Three of the four are empty for most entries, and the emptiness is the most useful thing the table shows: speech is the least documented part of generative video.
Reading the table from left to right answers a practical question in order — can this model speak, can it speak as a particular character, can it speak this language, and can the lines be made to land where the edit needs them.
1The table
- Speech — Seventeen models on four speech fields: where audio comes from, which languages are named, how a voice attaches to a character, and what drives lip-sync.
A column is added when enough vendors document something to compare; one nobody addresses stays a question in the notes.
Other notes: Models, Fields, Routes, Side by side, Questions, Terms, Learn, Data. What counts as a documented control is set out on the reading page.