Vidu: three tracks in one pass, on by default
Q3 outputs speech, sound effects and music natively, on by default in the API, with a speech-only mode beside it. The default decides what most output from this model sounds like, because it applies to every call made without reading the parameter list. As of 2026-09-12.
| What the documentation settles | What it leaves to a take |
|---|---|
| Q3 outputs speech, sound effects and music natively, on by default in the API | How the three tracks balance against each other in a returned file |
| A speech-only mode returns the dialogue without the effects and music | Whether that mode changes the delivery as well as the mix |
| Nothing is published about where a voice comes from | Whether a voice can be chosen, supplied or held to a character at all |
Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.
1A default is a design decision disguised as a parameter
Audio on by default means the baseline output of this model is a finished soundscape rather than silent footage waiting for post. A team benchmarking it against a silent competitor is comparing two different deliverables, and the one that arrives with music will usually feel further along.
The same default makes silence something that has to be asked for. A scene meant to play dry will not unless the call says so, and the fix after the fact is a regeneration rather than a mute, because the bed is inside the file.
2Separating speech from the bed is the useful half
A speech-only mode is the one documented affordance in this register for keeping the dialogue and replacing everything under it. For drama that is the mode that matters, because the music is a creative decision that belongs to an editor rather than to a generator.
What the reference does not carry is any statement about the voice itself. Three tracks are described and nobody is identified as speaking, which is enough for a single clip and not enough for a cast.
3The others that make sound while they make the picture
Seven more entries generate audio in the same pass. This is the only one that documents a way to keep the speech and drop the rest.
- Kling AI — with the picture, on video 3.0.
- LTX Studio — with the picture, and audio to video.
- MiniMax — with the picture.
- SceneMixer — with the picture, in the language set for the project.
- Sora 2 — with the picture.
- Veo — with the picture.
- Wan 3.0 — with the picture, unless switched off.
- Audio sourceQ3 outputs speech, sound effects and music natively, on by default in the API, with a speech-only modethree tracks in one pass
- Per-character bindingNot documented by the vendor (as of 2026-09-12)no public description
4Sources
Read from the API reference at platform.vidu.com on 2026-09-12. The same column across every entry is on audio source; everything this vendor publishes about speech is on Vidu. What counts as documented is on how read.