Wan 3.0 and Wan2.2-S2V: two opposite architectures
One defaults its audio parameter to true and returns a track unless told not to. The other cannot start without a recording, because the audio is the driving input. Same family name, opposite production shapes. As of 2026-09-22.
| Field | Wan 3.0 | Wan2.2-S2V | Where they part |
|---|---|---|---|
| Audio source | With the picture, unless switched off | Supplied; the model is audio-driven | Different answers |
| Voice source | Reference audio, 15 seconds in total | The audio track itself | Different answers |
| Per-character binding | On the call | Not documented by the vendor | Only Wan 3.0 answers |
| Languages | Not documented by the vendor | Not documented by the vendor | Same answer |
| Lip-sync | Not documented by the vendor | The audio input | Only Wan2.2-S2V answers |
Inclusion rule. Two entries are given a page together when at least one column puts them at opposite grades of answer. Pairs that agree on every column, or that are both blank throughout, do not get a page. Order. Fixed field order, identical on every side-by-side page.
1A family name says nothing about the architecture
These two arrive under the same brand and belong in different chapters of a production plan. One is a generator that speaks; the other is a stage that animates a face to speech somebody else recorded. A team that picked one on the strength of the name and built a schedule around it would have to rebuild the schedule to use the other.
That is a general warning rather than a complaint about this vendor. Model families grow sideways, and the question worth asking about any new member is which direction the audio flows rather than what the version number suggests.
2One publishes figures, the other publishes a licence
The hosted entry publishes a default, a pricing statement, a format, a clip window and a byte ceiling. The model card publishes two resolutions and a permissive licence. Those are two different kinds of commitment: one binds the vendor, the other hands the reader the means to check anything unbound.
So the blanks are not comparable either. An unstated behaviour on a hosted API stays unstated; an unstated behaviour on published weights is a measurement somebody has not taken yet.
3Where they meet is the mouth, and only one names a driver
The audio-driven card states that the result stays synchronised to the audio input, and restates it where a pose argument is introduced. The hosted reference publishes a dialogue syntax, a default and a price, and never mentions a mouth at all.
Read together they make the pattern in this register explicit: naming a driver is easy when the waveform is an input and apparently hard when it is an output, even for the same vendor.
4Each of them on its own
The column this pair was chosen for is audio source, and each entry has a page of its own on it. The full row for either, all five columns with the wording behind each cell, is on its model note.
- Wan 3.0 on audio source — with the picture, unless switched off.
- Wan2.2-S2V on audio source — supplied; the model is audio-driven.
- Wan 3.0, all five fields — read from the video generation API reference.
- Wan2.2-S2V, all five fields — read from the model card.
- Audio sourceThe audio parameter defaults to true so the returned video carries an audio track, and false returns one withouton unless it is switched off
- What choosing silence costsEnabling or disabling audio does not affect pricing
- Audio sourceThe card describes audio-driven cinematic video generation from an audio input with a reference image and an optional text prompthanded to the model rather than made by it
- A second control alongside the audioA pose video argument lets the result follow a pose sequence while staying synchronised to the audio
- What the card does settleThe card states support for 480P and 720P and licenses the weights under Apache 2.0
5Sources
Every cell above is read from the documentation each vendor publishes, on the dates carried by the two model notes. The whole register on one column is on the field note; entries grouped by the shape of their answer are on the routes. Other pairs: Veo and Hedra, Vidu and sync-3.