Sentioscope

Speech and voice controls, as each vendor documents them

Wan 3.0 and Wan2.2-S2V: two opposite architectures

One defaults its audio parameter to true and returns a track unless told not to. The other cannot start without a recording, because the audio is the driving input. Same family name, opposite production shapes. As of 2026-09-22.

One family name, and audio flowing two waysOne entry returns a track unless told not to; the other cannot start without a recording, because the audio is the driving input. The direction of that arrow decides the whole shape of a production plan.Wan 3.0Wan2.2-S2VAudio directionOut of the model, on by defaultInto the model, and requiredWhat is publishedDefault, price, format, clip window,bytesTwo resolutions and a licenceThe mouthNever mentionedStated to stay synchronised to the audioHow a blank readsUnstated on a hosted APIA measurement nobody has takenSame brand, different chapters of a plan
Fig. 1 The question worth asking about any new member of a model family is which way the audio flows, not what the version number suggests.
Controls documented, model by modelFilled where the model documents that control, hollow where nothing is published about it.Controls documented, model by modelWan 3.0Wan2.2-S2VWhere they partAudio sourceAudio source — Wan 3.0: With the picture, unless switched offAudio source — Wan2.2-S2V: Supplied; the model is audio-drivenAudio source — Where they part: Different answersVoice sourceVoice source — Wan 3.0: Reference audio, 15 seconds in totalVoice source — Wan2.2-S2V: The audio track itselfVoice source — Where they part: Different answersPer-character bindingPer-character binding — Wan 3.0: On the callPer-character binding — Wan2.2-S2V: Not documented by the vendorPer-character binding — Where they part: Only Wan 3.0 answersLanguagesLanguages — Wan 3.0: Not documented by the vendorLanguages — Wan2.2-S2V: Not documented by the vendorLanguages — Where they part: Same answerLip-syncLip-sync — Wan 3.0: Not documented by the vendorLip-sync — Wan2.2-S2V: The audio inputLip-sync — Where they part: Only Wan2.2-S2V answers
Fig. 2 Filled where the model documents that control, hollow where nothing is published about it.
Models documenting each controlHow many models document each control. A hollow column is a statement about documentation, not capability.Models documenting each controlWan 3.03 of 5Wan2.2-S2V3 of 5Where they part5 of 5
Fig. 3 How many models document each control. A hollow column is a statement about documentation, not capability.
Wan 3.0 and Wan2.2-S2V on the five fields this register keeps, with the gap on each. Recorded 2026-09-22.
FieldWan 3.0Wan2.2-S2VWhere they part
Audio sourceWith the picture, unless switched offSupplied; the model is audio-drivenDifferent answers
Voice sourceReference audio, 15 seconds in totalThe audio track itselfDifferent answers
Per-character bindingOn the callNot documented by the vendorOnly Wan 3.0 answers
LanguagesNot documented by the vendorNot documented by the vendorSame answer
Lip-syncNot documented by the vendorThe audio inputOnly Wan2.2-S2V answers

Inclusion rule. Two entries are given a page together when at least one column puts them at opposite grades of answer. Pairs that agree on every column, or that are both blank throughout, do not get a page. Order. Fixed field order, identical on every side-by-side page.

1A family name says nothing about the architecture

These two arrive under the same brand and belong in different chapters of a production plan. One is a generator that speaks; the other is a stage that animates a face to speech somebody else recorded. A team that picked one on the strength of the name and built a schedule around it would have to rebuild the schedule to use the other.

That is a general warning rather than a complaint about this vendor. Model families grow sideways, and the question worth asking about any new member is which direction the audio flows rather than what the version number suggests.

2One publishes figures, the other publishes a licence

The hosted entry publishes a default, a pricing statement, a format, a clip window and a byte ceiling. The model card publishes two resolutions and a permissive licence. Those are two different kinds of commitment: one binds the vendor, the other hands the reader the means to check anything unbound.

So the blanks are not comparable either. An unstated behaviour on a hosted API stays unstated; an unstated behaviour on published weights is a measurement somebody has not taken yet.

3Where they meet is the mouth, and only one names a driver

The audio-driven card states that the result stays synchronised to the audio input, and restates it where a pose argument is introduced. The hosted reference publishes a dialogue syntax, a default and a price, and never mentions a mouth at all.

Read together they make the pattern in this register explicit: naming a driver is easy when the waveform is an input and apparently hard when it is an output, even for the same vendor.

4Each of them on its own

The column this pair was chosen for is audio source, and each entry has a page of its own on it. The full row for either, all five columns with the wording behind each cell, is on its model note.

  • Audio source
    The audio parameter defaults to true so the returned video carries an audio track, and false returns one withouton unless it is switched offAlibaba Cloud, Wan 3.0 API reference / recorded 2026-09-22
  • What choosing silence costs
    Enabling or disabling audio does not affect pricingAlibaba Cloud, Wan 3.0 API reference / recorded 2026-09-22
  • Audio source
    The card describes audio-driven cinematic video generation from an audio input with a reference image and an optional text prompthanded to the model rather than made by itWan-AI, Wan2.2-S2V-14B model card / recorded 2026-09-22
  • A second control alongside the audio
    A pose video argument lets the result follow a pose sequence while staying synchronised to the audioWan-AI, Wan2.2-S2V-14B model card / recorded 2026-09-22
  • What the card does settle
    The card states support for 480P and 720P and licenses the weights under Apache 2.0Wan-AI, Wan2.2-S2V-14B model card / recorded 2026-09-22

5Sources

Every cell above is read from the documentation each vendor publishes, on the dates carried by the two model notes. The whole register on one column is on the field note; entries grouped by the shape of their answer are on the routes. Other pairs: Veo and Hedra, Vidu and sync-3.