What kind of document each answer was read from
Answers in this register come from four kinds of document, and the kind predicts the shape. References publish parameters. Model cards publish capabilities and licences. Prompting guides publish conventions. Checklists publish restrictions. As of 2026-09-12.
| Model | What is documented | Detail |
|---|---|---|
| An API reference | Parameters, defaults, formats and ceilings | Nothing about what a result looks like |
| A model card | Capabilities, resolutions and a licence | Nothing a caller sets |
| A prompting guide | Where to write things, and rough budgets | Nothing enforceable |
| A compliance checklist | What a production may not do | Nothing about how anything works |
Inclusion rule. Every kind of vendor document read for the entries here is listed, with what that kind publishes and what it leaves out. A kind is listed once however many entries it covers. Order. Fixed order, from the most common kind to the least.
1A parameter list has no field for a convincing mouth
Most blanks in the lip-sync column come from API references, and the reason is structural: those documents describe what a caller sets, and nothing a caller sets makes a mouth look right. The subject falls outside what the page is for.
That explains the gap and does not close it. A producer choosing between tools before any money moves has only the documents, and a document written for a developer will not answer a question a producer has.
2A licence changes what a blank is worth
On a hosted API an absent sentence is a refusal to commit, and stays absent. On a model card with published weights the same absence is a measurement somebody has not taken, because anyone can run the thing and find out.
This register records the card rather than the code, because a comparison only works if every entry is read the same way. The licence is logged as its own fact so a reader can discount those blanks accordingly.
3A prompting guide is the kind that speaks to a writer
One entry's guide says where a line goes, how to label two speakers and roughly how much speech fits four seconds. None of that is enforceable and all of it is actionable, which is the reverse of everything an API reference offers.
It is also the kind of document most likely to go stale quietly, because nothing breaks when a convention changes. A production copying a guide into its own shot list should record the date it copied.
- Audio sourceThe audio parameter defaults to true so the returned video carries an audio track, and false returns one withouton unless it is switched off
- What the card does settleThe card states support for 480P and 720P and licenses the weights under Apache 2.0
- How a spoken line is written into a promptDialogue belongs in a dialogue block below the prose description, with exchanges limited to a handful of sentences
- What the vendor rules outCloning an identifiable voice is described as a likeness issue, with US digital-replica statutes named
4Sources
Every line is read from the vendor documentation linked on the speech controls page, checked 2026-09-12. What counts as documented is on how read. Nearby: Formats and file sizes, What sets the length, Charging for audio.