Sentioscope

Speech and voice controls, as each vendor documents them

Wan 3.0: caps in writing, and no language anywhere

This is the most precisely documented audio section in the register. Defaults, a pricing statement, clip lengths, file sizes and the exact place a spoken line goes. The language that line is delivered in is not mentioned once. As of 2026-09-22.

Megabytes in writing, and no language anywhereThis reference publishes a default, a pricing consequence, clip lengths, a byte ceiling and the exact place a spoken line goes. The language that line is delivered in is the one thing it never mentions.Published preciselyNever addressedThe switchDefaults to true; false returns no trackWhat language the default speaksThe clip1 to 15 s, 15 s in total, up to 15 MBWhether the clip sets the languageThe lineWritten into the prompt, after sayingWhether the script it is written in isusedThe priceIdentical with audio on or offNothing about markets at allSpace and inclination were both there
Fig. 1 A production could reasonably infer that the reference clip decides it, and an inference has no page to point at when behaviour changes.
Wan 3.0 on languages, statement by statement. Read from the vendor's video generation API reference on 2026-09-22.
What the documentation settlesWhat it leaves to a take
The audio parameter defaults to true; false returns a file with no audio trackWhat language the default track speaks when nothing is specified
Reference audio is WAV or MP3, one to fifteen seconds a clip, fifteen in total, up to 15 MBWhether the reference clip decides the language of the read
The worked example puts the spoken line in the prompt, after the word sayingWhether a line written in another script is performed in it

Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.

1Precision in one place is not precision everywhere

An API reference that publishes megabyte limits and a pricing consequence is not a vague document. That makes this blank harder to explain than the blanks on pages that never discuss sound: the vendor clearly had the space and the inclination to be specific, and this particular question did not come up.

The likeliest reading is that language is treated as an emergent property of the prompt and the reference clip rather than as a parameter. That would be a coherent design and it is not what the page says, so the cell stays unstated.

2A reference clip probably decides it, which is not the same as documenting it

Where a voice is established from a supplied recording, the language of that recording is the obvious lever. A production could reasonably plan on it. But planning on an inference means a change in the model's behaviour arrives with no warning and no page to point at.

The prompt syntax adds a second inference. If the line goes into the prompt after the word saying, then the script the line is written in is available to the model. Whether it is used that way is exactly the kind of thing a sentence could settle and does not.

3The other entries that never say which language

Eight more entries reach this column empty. Most of them say much less about audio than this one does, which is what makes the omission here worth a page.

  • LTX Studio — not documented by the vendor.
  • Luma Ray — not documented by the vendor.
  • MiniMax — not documented by the vendor.
  • Runway — not documented by the vendor.
  • Sora 2 — not documented by the vendor.
  • Veo — not documented by the vendor.
  • Vidu — not documented by the vendor.
  • Wan2.2-S2V — not documented by the vendor.
  • Voice source
    Reference audio is WAV or MP3, one to fifteen seconds a clip, fifteen seconds in total and no more than 15 MBa published capAlibaba Cloud, Wan 3.0 API reference / recorded 2026-09-22
  • How a spoken line is written into a prompt
    The worked example puts the spoken line in the prompt itself, after the word sayingAlibaba Cloud, Wan 3.0 API reference / recorded 2026-09-22
  • Audio source
    The audio parameter defaults to true so the returned video carries an audio track, and false returns one withouton unless it is switched offAlibaba Cloud, Wan 3.0 API reference / recorded 2026-09-22

4Sources

Read from the video generation API reference at alibabacloud.com on 2026-09-22. The same column across every entry is on languages; everything this vendor publishes about speech is on Wan 3.0. What counts as documented is on how read.