Wan 3.0: fifteen seconds, WAV or MP3, under 15 MB
The reference publishes every limit a production would ask for: WAV or MP3, one to fifteen seconds a clip, fifteen seconds in total, and no more than fifteen megabytes. Format, duration and size in one sentence. As of 2026-09-22.
| What the documentation settles | What it leaves to a take |
|---|---|
| Reference audio is WAV or MP3, one to fifteen seconds a clip, fifteen in total, up to 15 MB | Whether fifteen seconds carries an accent as well as a timbre |
| The audio parameter defaults to true; false returns a file with no audio track | What voice the default track uses when no reference is supplied |
| The worked example puts the spoken line in the prompt, after the word saying | How a reference voice and a written line interact |
Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.
1Publishing the format is rarer than publishing the duration
Several entries here give a length. This one also gives the container and the byte ceiling, which removes an entire class of failed call from a production's first day. WAV or MP3, nothing else implied, and a size that a phone recording will not exceed.
The completeness is worth noting because it is not matched elsewhere on the same page. This reference is precise about what a voice can be built from and silent about which language the resulting voice speaks.
2An identical total to another entry, reached differently
Fifteen seconds in total is the same figure another entry in this column publishes, and the surrounding arrangement is not the same. Here a clip may run the full fifteen seconds on its own; there the total is divided across at most three clips.
The difference decides how a cast is supplied. One long clean take against several short ones is a real choice, and it is the kind of detail that only becomes visible when two references are read against each other rather than in isolation.
3The others that take a clip on the call that uses it
Two more entries establish a voice from audio supplied with each generation. One publishes the same fifteen-second total split across three clips; the other takes the entire performance as its input.
- MiniMax — a reference clip on the call.
- Wan2.2-S2V — the audio track itself.
- Voice sourceReference audio is WAV or MP3, one to fifteen seconds a clip, fifteen seconds in total and no more than 15 MBa published cap
- Audio sourceThe audio parameter defaults to true so the returned video carries an audio track, and false returns one withouton unless it is switched off
- How a spoken line is written into a promptThe worked example puts the spoken line in the prompt itself, after the word saying
4Sources
Read from the video generation API reference at alibabacloud.com on 2026-09-22. The same column across every entry is on voice source; everything this vendor publishes about speech is on Wan 3.0. What counts as documented is on how read.