What Wan 3.0 documents about speech
Wan 3.0 (alibabacloud.com) returns audio unless a caller switches it off, states that the switch does not change the price, caps reference audio at fifteen seconds, and puts the spoken line inside the prompt after the word saying. As of 2026-09-22.
| Field | What the vendor documents |
|---|---|
| Audio source | On by default; false returns a file with no audio track |
| Languages | Not documented |
| Voice source | Reference audio, 15 seconds in total, WAV or MP3 up to 15 MB |
| Per-character binding | By reference audio supplied on the call |
| Lip-sync | Not documented |
Inclusion rule. Fields are the same five for every model here; a field the vendor does not document is recorded as not documented rather than inferred from generated output. Order. Fixed field order, identical on every model page.
1A default with a price attached to it
The audio flag defaults to true, so the ordinary return from this model is a file with sound in it. The reference then does something almost nobody does: it says that enabling or disabling audio does not affect pricing. Silence costs what speech costs.
That turns a toggle into a production decision. A team rendering mute picture for a composer to score is paying for a soundtrack it discards, and knowing that in advance is worth more than a paragraph of capability copy.
2Fifteen seconds, in writing
Reference audio is WAV or MP3, one to fifteen seconds a clip, fifteen seconds in total, no larger than fifteen megabytes. The total is the number that binds: a scene with three speakers splits one fifteen-second budget between them, so voices tend to be supplied a shot at a time.
Because the clip travels on the call rather than being stored, the clip becomes a production asset in its own right. Lose it and a character's voice is gone, and nothing in a returned file records which sample produced it.
3The line goes in beside the camera direction
The worked example puts the spoken words in the prompt itself, introduced by saying, in the same breath as the description of what is in frame. Dialogue is not a field to be filled but a clause in the paragraph that also carries the blocking.
That syntax has to be written for, and it does not travel: another entry here wants the line in a block of its own, and a third takes a recording instead of words. The language of the performance is not addressed anywhere on the reference, which leaves the commonest question about a spoken line unanswered.
4This entry, one column at a time
Each of these stays inside a single field: what Wan 3.0 puts there, what the wording settles, and what it leaves for a take to answer.
- Audio source — with the picture, unless switched off.
- Voice source — reference audio, 15 seconds in total.
- Per-character binding — on the call.
- Languages — not documented by the vendor.
- Lip-sync — not documented by the vendor.
5Read against another entry
Each of these puts Wan 3.0 beside one other entry on a column where the two land at opposite grades of answer.
- Wan 3.0 and Wan2.2-S2V — on audio source.
- MiniMax and Wan 3.0 — on voice source.
- Sora 2 and Wan 3.0 — on audio source.
6Where this model sits against the rest
| Model | Audio source |
|---|---|
| D-ID | A text script read aloud, or an audio url supplied |
| Hedra | Supplied; audio is a required input |
| HeyGen | A text to speech endpoint, apart from the picture |
| Kling AI | Generated with the picture, VIDEO 3.0 |
| LTX Studio | Generated jointly with the picture, plus audio-to-video |
| Luma Ray | Not documented; the video reference does not mention sound |
| MiniMax | Native speech, generated with the picture |
| PixVerse | Supplied to a speech endpoint, or read from text there |
| Runway | Not documented where voices are defined |
| SceneMixer | Generated with the picture, in the language set for the project |
| Sora 2 | Generated with the picture; video and audio are both listed as output |
| sync-3 | Supplied, or read from text on the call |
| Synthesia | Synthesised from the script, or an uploaded recording |
| Veo | Generated with the video rather than added afterwards |
| Vidu | Q3 generates speech, effects and music natively, on by default |
| Wan 3.0 | On by default; false returns a file with no audio track |
| Wan2.2-S2V | Supplied; the card describes the model as audio-driven |
Inclusion rule. Models whose vendor documents something about generated speech. A model with nothing published on the point is still listed, with the cell marked hollow. Order. Alphabetical by model name.
| Model | Languages |
|---|---|
| D-ID | A language field beside the voice; no list on the reference |
| Hedra | Full multi-language support claimed, none named |
| HeyGen | Thirty or more, counted on the translation endpoint |
| Kling AI | Five, not named |
| LTX Studio | Not documented |
| Luma Ray | Not documented |
| MiniMax | Not documented by the vendor |
| PixVerse | Multiple, with speech, singing and advertisements named as types |
| Runway | Not documented |
| SceneMixer | 15, named, with Cantonese a 16th for dialogue only |
| Sora 2 | Not documented |
| sync-3 | Ninety-five or more, counted rather than named |
| Synthesia | A catalogue, each voice carrying a language name and code |
| Veo | Not documented |
| Vidu | Not documented |
| Wan 3.0 | Not documented |
| Wan2.2-S2V | Not documented |
Inclusion rule. Models whose vendor documents something about generated speech. A model with nothing published on the point is still listed, with the cell marked hollow. Order. Alphabetical by model name.
| Model | Voice source |
|---|---|
| D-ID | A voice id from one of five named speech providers |
| Hedra | An uploaded track, or a voice id from the voices endpoint |
| HeyGen | A clone, instant from one recording or professional from 20 minutes or more |
| Kling AI | Not documented by the vendor |
| LTX Studio | Not documented |
| Luma Ray | Not documented |
| MiniMax | Reference audio, 15 s total across 3 clips |
| PixVerse | Built-in voices, or a custom voice from a supplied sample |
| Runway | An audio sample of 10 seconds to 5 minutes, or a description in words |
| SceneMixer | Preset voice library, or the user's own recordings |
| Sora 2 | Not documented |
| sync-3 | Not documented on the model page |
| Synthesia | The catalogue, or a cloned voice with a language list of its own |
| Veo | Not documented |
| Vidu | Not documented |
| Wan 3.0 | Reference audio, 15 seconds in total, WAV or MP3 up to 15 MB |
| Wan2.2-S2V | Whatever track is handed over; no catalogue |
Inclusion rule. Models whose vendor documents something about generated speech. A model with nothing published on the point is still listed, with the cell marked hollow. Order. Alphabetical by model name.
7Sources
Taken from the video generation API reference at alibabacloud.com on 2026-09-22. Every entry against the same five fields sits on the speech controls page; the rule behind each field is on how read. The entries either side of this one: Wan2.2-S2V, D-ID.
- Audio sourceThe audio parameter defaults to true so the returned video carries an audio track, and false returns one withouton unless it is switched off
- What choosing silence costsEnabling or disabling audio does not affect pricing
- Voice sourceReference audio is WAV or MP3, one to fifteen seconds a clip, fifteen seconds in total and no more than 15 MBa published cap
- How a spoken line is written into a promptThe worked example puts the spoken line in the prompt itself, after the word saying