Wan 3.0: on unless switched off, and priced the same
The audio parameter defaults to true, so a returned video carries a track unless the call says false. The reference adds the part that changes a budget conversation: enabling or disabling audio does not affect pricing. As of 2026-09-22.
| What the documentation settles | What it leaves to a take |
|---|---|
| The audio parameter defaults to true; false returns a file with no audio track | What the default track contains when the prompt asks for nothing |
| Enabling or disabling audio does not affect pricing | Whether the generated audio is good enough to keep at that rate |
| The worked example puts the spoken line in the prompt, after the word saying | How several speakers in one shot are written in that syntax |
Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.
1Silence costs the same, which makes it a creative decision
Most audio features are sold as an upgrade, so a team that does not want the sound expects to save by switching it off. Here the vendor says plainly that the rate does not move. Generating silent picture therefore means paying for a capability and discarding it, and the only reason left to do so is that the sound is not wanted.
Turned the other way, the same sentence says the audio is free to try on every shot. Whether it is good enough to keep is then the whole question, and the cost of replacing it lands entirely outside the rate card.
2The line goes in the prompt, beside the camera direction
Putting the spoken words after the word saying, inside the same prose that describes the shot, is a syntax rather than a parameter. It means the dialogue and the staging are edited in the same string, and a shot list has to carry them together or convert them at generation time.
That is the part of a shot list that travels least well between models. One vendor wants a labelled block, another takes reference voices, this one wants a clause. Keeping the line in plain language alongside whichever syntax is current is the habit that survives a model change.
3The others that make sound while they make the picture
Seven more entries generate audio in the same pass. This is the only one that states what choosing silence costs.
- Kling AI — with the picture, on video 3.0.
- LTX Studio — with the picture, and audio to video.
- MiniMax — with the picture.
- SceneMixer — with the picture, in the language set for the project.
- Sora 2 — with the picture.
- Veo — with the picture.
- Vidu — with the picture, speech available on its own.
- Audio sourceThe audio parameter defaults to true so the returned video carries an audio track, and false returns one withouton unless it is switched off
- What choosing silence costsEnabling or disabling audio does not affect pricing
- How a spoken line is written into a promptThe worked example puts the spoken line in the prompt itself, after the word saying
4Sources
Read from the video generation API reference at alibabacloud.com on 2026-09-22. The same column across every entry is on audio source; everything this vendor publishes about speech is on Wan 3.0. What counts as documented is on how read.