Wan 3.0: caps in writing, and no language anywhere
This is the most precisely documented audio section in the register. Defaults, a pricing statement, clip lengths, file sizes and the exact place a spoken line goes. The language that line is delivered in is not mentioned once. As of 2026-09-22.
| What the documentation settles | What it leaves to a take |
|---|---|
| The audio parameter defaults to true; false returns a file with no audio track | What language the default track speaks when nothing is specified |
| Reference audio is WAV or MP3, one to fifteen seconds a clip, fifteen in total, up to 15 MB | Whether the reference clip decides the language of the read |
| The worked example puts the spoken line in the prompt, after the word saying | Whether a line written in another script is performed in it |
Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.
1Precision in one place is not precision everywhere
An API reference that publishes megabyte limits and a pricing consequence is not a vague document. That makes this blank harder to explain than the blanks on pages that never discuss sound: the vendor clearly had the space and the inclination to be specific, and this particular question did not come up.
The likeliest reading is that language is treated as an emergent property of the prompt and the reference clip rather than as a parameter. That would be a coherent design and it is not what the page says, so the cell stays unstated.
2A reference clip probably decides it, which is not the same as documenting it
Where a voice is established from a supplied recording, the language of that recording is the obvious lever. A production could reasonably plan on it. But planning on an inference means a change in the model's behaviour arrives with no warning and no page to point at.
The prompt syntax adds a second inference. If the line goes into the prompt after the word saying, then the script the line is written in is available to the model. Whether it is used that way is exactly the kind of thing a sentence could settle and does not.
3The other entries that never say which language
Eight more entries reach this column empty. Most of them say much less about audio than this one does, which is what makes the omission here worth a page.
- LTX Studio — not documented by the vendor.
- Luma Ray — not documented by the vendor.
- MiniMax — not documented by the vendor.
- Runway — not documented by the vendor.
- Sora 2 — not documented by the vendor.
- Veo — not documented by the vendor.
- Vidu — not documented by the vendor.
- Wan2.2-S2V — not documented by the vendor.
- Voice sourceReference audio is WAV or MP3, one to fifteen seconds a clip, fifteen seconds in total and no more than 15 MBa published cap
- How a spoken line is written into a promptThe worked example puts the spoken line in the prompt itself, after the word saying
- Audio sourceThe audio parameter defaults to true so the returned video carries an audio track, and false returns one withouton unless it is switched off
4Sources
Read from the video generation API reference at alibabacloud.com on 2026-09-22. The same column across every entry is on languages; everything this vendor publishes about speech is on Wan 3.0. What counts as documented is on how read.