Sentioscope

Speech and voice controls, as each vendor documents them

Wan 3.0: fifteen seconds, WAV or MP3, under 15 MB

The reference publishes every limit a production would ask for: WAV or MP3, one to fifteen seconds a clip, fifteen seconds in total, and no more than fifteen megabytes. Format, duration and size in one sentence. As of 2026-09-22.

Format, duration and size, all in one sentenceWAV or MP3, one to fifteen seconds a clip, fifteen seconds in total and no more than fifteen megabytes. Publishing the container alongside the duration removes a class of failed call from a production's first day.Everything a production would have had to askFormatWAV or MP3, with nothing else implied or left to discovery.Per clipOne to fifteen seconds, so a single take may fill the wholebudget.In totalFifteen seconds for the call, however many clips it is splitacross.SizeNo more than fifteen megabytes, which a phone recording will notexceed.And no statement about which language it speaks
Fig. 1 The same fifteen-second total appears elsewhere in this column split across three clips, which is a real choice about how a cast is supplied.
Wan 3.0 on voice source, statement by statement. Read from the vendor's video generation API reference on 2026-09-22.
What the documentation settlesWhat it leaves to a take
Reference audio is WAV or MP3, one to fifteen seconds a clip, fifteen in total, up to 15 MBWhether fifteen seconds carries an accent as well as a timbre
The audio parameter defaults to true; false returns a file with no audio trackWhat voice the default track uses when no reference is supplied
The worked example puts the spoken line in the prompt, after the word sayingHow a reference voice and a written line interact

Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.

1Publishing the format is rarer than publishing the duration

Several entries here give a length. This one also gives the container and the byte ceiling, which removes an entire class of failed call from a production's first day. WAV or MP3, nothing else implied, and a size that a phone recording will not exceed.

The completeness is worth noting because it is not matched elsewhere on the same page. This reference is precise about what a voice can be built from and silent about which language the resulting voice speaks.

2An identical total to another entry, reached differently

Fifteen seconds in total is the same figure another entry in this column publishes, and the surrounding arrangement is not the same. Here a clip may run the full fifteen seconds on its own; there the total is divided across at most three clips.

The difference decides how a cast is supplied. One long clean take against several short ones is a real choice, and it is the kind of detail that only becomes visible when two references are read against each other rather than in isolation.

3The others that take a clip on the call that uses it

Two more entries establish a voice from audio supplied with each generation. One publishes the same fifteen-second total split across three clips; the other takes the entire performance as its input.

  • Voice source
    Reference audio is WAV or MP3, one to fifteen seconds a clip, fifteen seconds in total and no more than 15 MBa published capAlibaba Cloud, Wan 3.0 API reference / recorded 2026-09-22
  • Audio source
    The audio parameter defaults to true so the returned video carries an audio track, and false returns one withouton unless it is switched offAlibaba Cloud, Wan 3.0 API reference / recorded 2026-09-22
  • How a spoken line is written into a prompt
    The worked example puts the spoken line in the prompt itself, after the word sayingAlibaba Cloud, Wan 3.0 API reference / recorded 2026-09-22

4Sources

Read from the video generation API reference at alibabacloud.com on 2026-09-22. The same column across every entry is on voice source; everything this vendor publishes about speech is on Wan 3.0. What counts as documented is on how read.