Whether the dialogue can be kept without the music
The usual objection to generated audio is the music rather than the speech. One entry documents a mode returning speech alone. Everywhere else the tracks arrive mixed into one file, and an editor cannot separate them afterwards. As of 2026-09-12.
| Model | What is documented | Detail |
|---|---|---|
| sync-3 | Whatever was supplied | Separation is upstream, not here |
| Vidu | Speech, effects and music, with a speech-only mode | On by default in the API |
| Wan 3.0 | One track, on unless switched off | No separation documented |
Inclusion rule. Entries whose documentation describes what the generated audio track contains, or a mode that changes it. An entry silent about the contents of its audio does not earn a row. Order. Alphabetical by model name.
1Music is a decision that belongs to an editor
A generated bed arrives with a creative choice already made, and it is a choice nobody asked for. For a single clip that is convenient; across a series it means the score is whatever the prompt produced on each shot, which is the opposite of how a score works.
A documented mode that keeps the dialogue and drops the rest removes the objection without removing the feature. Only one entry here publishes one, and it is the entry whose default is the fullest mix.
2Baked-in audio makes a mute impossible
Where the tracks are mixed into the returned file, an unwanted bed cannot be muted in the edit; the shot has to be generated again with different instructions. That converts a small aesthetic objection into a cost and a delay.
It also means a scene meant to play dry needs the prompt to say so. One vendor documents the phrasing for asking for silence, which is unusual and worth copying: an explicit negative rather than an absence.
3Supplied-audio products never have the problem
Where the sound is an input, whatever was handed over is what plays, so separation happened upstream in a mix a production controls. That is a real advantage of the audio-driven arrangement and it is rarely stated as one.
The cost, as always with that arrangement, is that the sound has to exist first. A team with a mix already has the cleanest path; a team hoping to hear something within the hour does not.
- Audio sourceQ3 outputs speech, sound effects and music natively, on by default in the API, with a speech-only modethree tracks in one pass
- Audio sourceThe audio parameter defaults to true so the returned video carries an audio track, and false returns one withouton unless it is switched off
- Audio sourceThe accepted pairs are video with audio, video with text, image with audio and image with textsupplied or read from text
4Sources
Every line is read from the vendor documentation linked on the speech controls page, checked 2026-09-12. What counts as documented is on how read. Nearby: Regenerating a shot, Consent in the docs, A second control.