Speech and voice controls, as each vendor documents them
What Vidu documents about speech
Vidu (platform.vidu.com) documents native audio on Q3 covering speech, sound effects and music, switched on by default in the API, with a mode that returns speech alone. As of 2026-09-12.
Fig. 1 Filled where the model documents that control, hollow where nothing is published about it.
The five fields this register keeps, for Vidu. Read from the vendor's API reference on 2026-09-12.
Inclusion rule. Fields are the same five for every model here; a field the vendor does not document is recorded as not documented rather than inferred from generated output. Order. Fixed field order, identical on every model page.
1A default that decides what most output sounds like
Audio being on by default is a small sentence with a large effect. Every call made without reading the parameter list returns a clip with sound, which means the baseline output of this model is a finished soundscape rather than a silent picture waiting for post. A team measuring the model against a silent competitor is not comparing like with like.
The speech-only mode is the counterpart and the more useful one for drama: it separates the line from the effects and music, so dialogue can be kept while the bed is replaced. No other model in this register documents that separation.
Fig. 2 The same reference page that carries this model's audio statement, written for a developer calling the endpoint rather than for a production.Screenshot of platform.vidu.com/docs/reference-to-video, captured 2026-09-16. Reproduced to show where this entry was read from.
Four of the five fields this register keeps are empty here. No language list or count, no statement about where a voice comes from, no way described to hold a voice to a character, and nothing on lip-sync. The API reference describes what is produced and not who is speaking.
For a single clip that is enough. For a series it leaves the central question open, because a cast that is re-drawn on every call is not a cast. This site records the gap rather than filling it from listening tests, which would be an opinion about output rather than a published fact.
Fig. 3 One field at a time, every entry side by side, with Vidu marked. The full wording for every cell is in the rows below. This entry was read 2026-09-12.
Every entry in this register on the same field, with Vidu marked. Each cell carries the reading date of its own entry; this one was read 2026-09-12.
Supplied; the card describes the model as audio-driven
Inclusion rule. Models whose vendor documents something about generated speech. A model with nothing published on the point is still listed, with the cell marked hollow. Order. Alphabetical by model name.
Fig. 4 One field at a time, every entry side by side, with Vidu marked. The full wording for every cell is in the rows below. This entry was read 2026-09-12.
Every entry in this register on the same field, with Vidu marked. Each cell carries the reading date of its own entry; this one was read 2026-09-12.
Inclusion rule. Models whose vendor documents something about generated speech. A model with nothing published on the point is still listed, with the cell marked hollow. Order. Alphabetical by model name.
Fig. 5 One field at a time, every entry side by side, with Vidu marked. The full wording for every cell is in the rows below. This entry was read 2026-09-12.
Every entry in this register on the same field, with Vidu marked. Each cell carries the reading date of its own entry; this one was read 2026-09-12.
Inclusion rule. Models whose vendor documents something about generated speech. A model with nothing published on the point is still listed, with the cell marked hollow. Order. Alphabetical by model name.
Taken from the API reference at platform.vidu.com on 2026-09-12. Every entry against the same five fields sits on the speech controls page; the rule behind each field is on how read. The entries either side of this one: Wan 3.0, Wan2.2-S2V.
Audio source
Q3 outputs speech, sound effects and music natively, on by default in the API, with a speech-only modethree tracks in one passVidu, API reference/ recorded 2026-09-12
Per-character binding
Not documented by the vendor (as of 2026-09-12)no public descriptionVidu, API reference/ recorded 2026-09-12