The speech notes as one file
The notes are exported as one file. Each row records what a vendor documents about generated speech, with the page it was read from. As of 2026-09-12.
Speech is the least documented part of generative video, so this file is unusually full of rows recording that nothing is published. That imbalance is the most informative thing in it.
Each row names the field it fills, so a reader rebuilding the comparison cannot count a reference-audio claim as a native-audio one. Those two get merged constantly in product copy and the merge changes what a production has to plan for.
Nothing here comes from listening. Every row is a documented statement, which means the file describes what vendors commit to rather than what their models sound like.
Reuse under the licence below is permitted, including commercially, provided the source and date columns stay attached to the rows.
1What each column holds
| Column | What it holds |
|---|---|
| fact_id | The row's handle. Rewriting a note does not change it. |
| value | What the model's documentation states about speech. |
| applies_to | The control or language list the row is about. |
| source_url | The documentation page it came from. |
| source_name | The vendor's own name for that page. |
| value_since | From when the documentation has said this. |
| verified | The last day that page was read. |
Inclusion rule. Every recorded field is exported, including the many that record no published control. Order. File order follows recording order, grouping rows by model.
The file is at /datasets/controls.csv, licensed CC BY 4.0. Attribution should name the notes and the reading date.
Other notes: Speech controls, Models, Fields, Routes, Side by side, Questions, Terms, Learn. What counts as a documented control is set out on the reading page.