Wan2.2-S2V: the result stays synchronised to the audio
Synchronisation to the audio input is the card's central claim, and it is restated where the pose argument is introduced: the result follows a pose sequence while staying synchronised to the audio. Two channels, one of them authoritative. As of 2026-09-22.
| What the documentation settles | What it leaves to a take |
|---|---|
| The card describes audio-driven video generation from an audio input with a reference image | How the mouth behaves when the recording holds two voices |
| A pose video argument lets the result follow a pose sequence while staying synchronised | What happens when the pose turns the face away from camera |
| Support is stated for 480P and 720P, with weights published under Apache 2.0 | Whether alignment is equally good at both resolutions |
Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.
1Restating synchronisation next to the pose argument is deliberate
A vendor introducing a second control has a choice about whether to say it does not break the first. This card says it: the pose sequence is followed while synchronisation holds. That is a precedence claim, and it is the only one in this column.
What it leaves open is the case where the two genuinely conflict. A pose that turns the head away while a line plays cannot both be followed and stay aligned, and the card does not say which gives way.
2Open weights turn an unstated detail into a measurement
On a hosted product the questions this cell leaves open would stay open. Here the weights are published under a permissive licence, so a team that needs to know how alignment behaves on overlapping speech or at the lower resolution can find out by running it.
The register still records the card rather than the code, because every entry has to be read the same way for the comparison to mean anything. The licence sits in the fact list so a reader can weigh this row's silences differently.
3The others that name what the mouth is following
Four more entries name a driver. This is the only one that also names a second control and states that it does not displace the first.
- Hedra — the supplied audio.
- MiniMax — whoever is on screen.
- sync-3 — the audio it is given.
- Synthesia — the spoken content, framed close.
- Audio sourceThe card describes audio-driven cinematic video generation from an audio input with a reference image and an optional text prompthanded to the model rather than made by it
- A second control alongside the audioA pose video argument lets the result follow a pose sequence while staying synchronised to the audio
- What the card does settleThe card states support for 480P and 720P and licenses the weights under Apache 2.0
4Sources
Read from the model card at huggingface.co on 2026-09-22. The same column across every entry is on lip-sync; everything this vendor publishes about speech is on Wan2.2-S2V. What counts as documented is on how read.