Wan 3.0: reference audio, supplied per generation
Reference audio is supplied per generation, inside a fifteen-second total. Nothing is stored, so a character's voice is only as consistent as the file a production reaches for on the fortieth call. As of 2026-09-22.
| What the documentation settles | What it leaves to a take |
|---|---|
| Reference audio is supplied per generation, WAV or MP3, fifteen seconds in total | Whether the same clip twice yields the same voice twice |
| The audio parameter defaults to true; false returns a file with no audio track | What voice the default track uses when no reference is supplied |
| The worked example puts the spoken line in the prompt, after the word saying | Whether the written line can specify who is speaking |
Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.
1A default track with no reference is an unowned voice
Audio is on unless it is switched off, and a reference clip is optional. That means the easiest call to make returns speech in a voice nobody chose, and the second-easiest returns it in a voice supplied for that call only. Neither is a cast.
The practical consequence is that consistency has to be opted into on every single generation. Forgetting the reference does not produce an error; it produces a different person saying the same line.
2Fifteen seconds is the whole cast budget for one call
With a fifteen-second total per generation, a scene with several speakers cannot carry all of them at once. Voices go in one at a time, which means the shot list rather than the platform is holding the mapping between characters and clips.
Filing discipline therefore substitutes for a binding mechanism. Name the clip after the character, keep it beside the visual reference, and never regenerate a shot with whichever file happens to be at hand.
3The others that re-establish the voice on each call
Four more entries put the voice on the request. Three pass a short identifier that cannot drift; this one and one other pass audio, which can.
- Voice sourceReference audio is WAV or MP3, one to fifteen seconds a clip, fifteen seconds in total and no more than 15 MBa published cap
- Audio sourceThe audio parameter defaults to true so the returned video carries an audio track, and false returns one withouton unless it is switched off
- How a spoken line is written into a promptThe worked example puts the spoken line in the prompt itself, after the word saying
4Sources
Read from the video generation API reference at alibabacloud.com on 2026-09-22. The same column across every entry is on per-character binding; everything this vendor publishes about speech is on Wan 3.0. What counts as documented is on how read.