Wan2.2-S2V: the voice is the file you hand over
There is no catalogue and no sample window, because the audio is not a reference for a voice but the performance itself. The card describes audio-driven generation from an audio input with a reference image and an optional prompt. As of 2026-09-22.
| What the documentation settles | What it leaves to a take |
|---|---|
| The card describes audio-driven video generation from an audio input with a reference image | How the result behaves when the recording carries two voices |
| Support is stated for 480P and 720P, with weights published under Apache 2.0 | What a self-hosted deployment inherits that the card omits |
| A pose video argument lets the result follow a pose sequence while staying synchronised | How a conflict between pose and audio resolves |
Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.
1A supplied performance is not a supplied reference
Every other entry taking audio in this column treats the clip as an example for the model to imitate. Here the clip is the delivery: whatever was recorded is what the audience hears, and the model's job is the picture. That removes the whole class of questions about how faithfully a voice is copied.
It also moves the casting entirely into a recording session, where the tools are microphones and performers rather than parameters. A production that already works this way loses nothing from an empty catalogue.
2Permissive weights change what an empty cell is worth
On a hosted API, an absent sentence is a refusal to commit. Here the weights are published under Apache 2.0, so a team that needs to know how the model treats a particular kind of recording can find out by running it rather than by asking.
The register still records the card, because a comparison only works if every entry is read the same way. The licence is logged as its own fact so a reader can discount this row's blanks accordingly.
3The others that take a clip on the call that uses it
Two more entries take audio with each generation. Both treat it as a reference with a fifteen-second ceiling; this one takes the finished performance and has no ceiling to publish.
- Audio sourceThe card describes audio-driven cinematic video generation from an audio input with a reference image and an optional text prompthanded to the model rather than made by it
- What the card does settleThe card states support for 480P and 720P and licenses the weights under Apache 2.0
- A second control alongside the audioA pose video argument lets the result follow a pose sequence while staying synchronised to the audio
4Sources
Read from the model card at huggingface.co on 2026-09-22. The same column across every entry is on voice source; everything this vendor publishes about speech is on Wan2.2-S2V. What counts as documented is on how read.