Which column every entry manages to fill
One column is almost always answered and the rest are mostly not. Where the sound comes from is filled by fifteen entries. Languages, a voice source, a persistence mechanism and a mouth driver are all minority answers. As of 2026-09-12.
| Model | What is documented | Detail |
|---|---|---|
| Audio source | Fifteen entries | The two blanks are a picture page and a voices page |
| Voice source | Ten entries | Blank where a voice arrives without being chosen |
| Per-character binding | Nine entries | Blank where nothing is said to outlast a call |
| Languages | Eight entries | And only two of the eight publish names |
| Lip-sync | Ten entries | Half of those name the feature and no driver |
Inclusion rule. Every column this register keeps is listed, with the number of entries whose documentation fills it. A column is not added unless enough vendors publish something comparable. Order. Fixed field order, the order the questions arrive in.
1The easiest question is the one everybody answers
Whether sound comes out of a model is a single binary fact, easy to state and hard to get wrong, so almost every document manages it. The columns that stay empty are the ones requiring a vendor to describe a workflow rather than a capability.
That pattern is the most consistent finding in this register. Documentation gets thinner exactly as the questions get more useful, and it thins in the same order across products that otherwise share nothing.
2The two blanks in the filled column are instructive
One is a video reference that never mentions sound; the other is a voices reference that never mentions a shot. Between them they cover both halves of this subject and neither covers the join, which is where every question here lives.
Neither is a weak product. The blanks describe what each document was written for, and a production is the party left to establish the rest by test.
3A filled cell still says less than it appears to
Counting filled cells would flatter the vendors who write most, so this register keeps the wording instead. Ten entries fill the mouth column and only five of those name anything as the driver; four name the feature and stop, and one argues the step away.
Grades inside a column matter more than the tally across it. A list beats a total, a driver beats a capability, and a stored voice beats a clip supplied on every call, regardless of how many cells are filled on either row.
- Audio sourceNot documented by the vendor (as of 2026-09-22)no mention of sound in the video reference
- Lip-syncNot documented by the vendor (as of 2026-09-22)not addressed where voices are defined
- Per-character bindingNot documented by the vendor (as of 2026-09-22)no public description
- LanguagesNot documented by the vendor (as of 2026-09-22)neither a list nor a count
- LanguagesNot documented by the vendor (as of 2026-09-22)nothing published about the spoken performance
4Sources
Every line is read from the vendor documentation linked on the speech controls page, checked 2026-09-12. What counts as documented is on how read. Nearby: Speech without the bed, Regenerating a shot, Consent in the docs.