Sentioscope

Speech and voice controls, as each vendor documents them

Wan 3.0: audio on by default, and mouths unaddressed

This reference is precise about almost everything else: the default, the price, the clip limits and where a spoken line goes in the prompt. Mouths are not mentioned, so the field is recorded as unstated. As of 2026-09-22.

Everything measured except the thing on screenA reference publishing megabyte ceilings and a pricing statement is not short of detail, which makes this blank different in kind from one on a page that never discusses sound at all.Published as a figureNever mentionedThe switchDefaults to true, false returns no trackWhat aligns the mouthThe clip1 to 15 s, up to 15 MB, WAV or MP3Whether the clip affects alignmentThe priceIdentical with audio on or offWhether two speakers are handledThe lineWritten into the prompt, after sayingWhether that helps the faceSpace and the habit of being specific were both there
Fig. 1 A production can infer that joint generation handles it, and an inference is worth less than a sentence when a shot comes back wrong.
Wan 3.0 on lip-sync, statement by statement. Read from the vendor's video generation API reference on 2026-09-22.
What the documentation settlesWhat it leaves to a take
The audio parameter defaults to true; false returns a file with no audio trackWhether the default track arrives aligned with the picture
The worked example puts the spoken line in the prompt, after the word sayingWhether writing the line that way improves alignment
Nothing in the reference addresses lip movementHow two speakers in one shot are handled

Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.

1Precision elsewhere makes this blank louder

A reference that publishes megabyte ceilings and a statement about pricing is not short of detail. That makes the absence here different in kind from the absence on a page that never discusses sound at all: the vendor had both the space and the habit of being specific.

The most likely reading is the same one that explains several blanks in this column. An API reference lists parameters, and there is no parameter for a convincing mouth, so the subject falls outside what the document is for.

2A dialogue syntax implies alignment without describing it

Telling a writer to put the spoken line in the prompt after the word saying means the model is expected to perform specific words. Something has to align a mouth to those words, and the reference does not name it.

A production can infer that the joint generation handles it, and an inference is worth less than a sentence when a shot comes back wrong. The register keeps the cell empty so that this row is comparable with the entries that did publish a driver.

3Others whose pages never raise the subject

Six more entries reach this column empty. This one publishes more detail about audio than most entries that filled it, which is what makes the omission worth a page.

  • D-ID — not documented by the vendor.
  • LTX Studio — not documented by the vendor.
  • Luma Ray — not documented by the vendor.
  • Runway — not documented by the vendor.
  • Veo — not documented by the vendor.
  • Vidu — not documented by the vendor.
  • Audio source
    The audio parameter defaults to true so the returned video carries an audio track, and false returns one withouton unless it is switched offAlibaba Cloud, Wan 3.0 API reference / recorded 2026-09-22
  • How a spoken line is written into a prompt
    The worked example puts the spoken line in the prompt itself, after the word sayingAlibaba Cloud, Wan 3.0 API reference / recorded 2026-09-22
  • Voice source
    Reference audio is WAV or MP3, one to fifteen seconds a clip, fifteen seconds in total and no more than 15 MBa published capAlibaba Cloud, Wan 3.0 API reference / recorded 2026-09-22

4Sources

Read from the video generation API reference at alibabacloud.com on 2026-09-22. The same column across every entry is on lip-sync; everything this vendor publishes about speech is on Wan 3.0. What counts as documented is on how read.