Sentioscope

Speech and voice controls, as each vendor documents them

Wan 3.0: reference audio, supplied per generation

Reference audio is supplied per generation, inside a fifteen-second total. Nothing is stored, so a character's voice is only as consistent as the file a production reaches for on the fortieth call. As of 2026-09-22.

Two easy calls, and neither of them is a castAudio is on unless switched off and a reference clip is optional, so the easiest call returns speech in a voice nobody chose and the next easiest returns it in a voice supplied for that call alone.What a Wan 3.0 call returns when audio is left onNo reference clipA voice nobody choseConsistent with nothing, and it will notbe flaggedA reference clipA voice for this callFifteen seconds in total, and theplatform stores none of itConsistency has to be opted into on every generation
Fig. 1 Forgetting the reference does not produce an error; it produces a different person saying the same line.
Wan 3.0 on per-character binding, statement by statement. Read from the vendor's video generation API reference on 2026-09-22.
What the documentation settlesWhat it leaves to a take
Reference audio is supplied per generation, WAV or MP3, fifteen seconds in totalWhether the same clip twice yields the same voice twice
The audio parameter defaults to true; false returns a file with no audio trackWhat voice the default track uses when no reference is supplied
The worked example puts the spoken line in the prompt, after the word sayingWhether the written line can specify who is speaking

Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.

1A default track with no reference is an unowned voice

Audio is on unless it is switched off, and a reference clip is optional. That means the easiest call to make returns speech in a voice nobody chose, and the second-easiest returns it in a voice supplied for that call only. Neither is a cast.

The practical consequence is that consistency has to be opted into on every single generation. Forgetting the reference does not produce an error; it produces a different person saying the same line.

2Fifteen seconds is the whole cast budget for one call

With a fifteen-second total per generation, a scene with several speakers cannot carry all of them at once. Voices go in one at a time, which means the shot list rather than the platform is holding the mapping between characters and clips.

Filing discipline therefore substitutes for a binding mechanism. Name the clip after the character, keep it beside the visual reference, and never regenerate a shot with whichever file happens to be at hand.

3The others that re-establish the voice on each call

Four more entries put the voice on the request. Three pass a short identifier that cannot drift; this one and one other pass audio, which can.

  • Voice source
    Reference audio is WAV or MP3, one to fifteen seconds a clip, fifteen seconds in total and no more than 15 MBa published capAlibaba Cloud, Wan 3.0 API reference / recorded 2026-09-22
  • Audio source
    The audio parameter defaults to true so the returned video carries an audio track, and false returns one withouton unless it is switched offAlibaba Cloud, Wan 3.0 API reference / recorded 2026-09-22
  • How a spoken line is written into a prompt
    The worked example puts the spoken line in the prompt itself, after the word sayingAlibaba Cloud, Wan 3.0 API reference / recorded 2026-09-22

4Sources

Read from the video generation API reference at alibabacloud.com on 2026-09-22. The same column across every entry is on per-character binding; everything this vendor publishes about speech is on Wan 3.0. What counts as documented is on how read.