Wan2.2-S2V: the language is whatever was recorded
Nothing on the card addresses spoken language, and for once the blank has a structural explanation rather than an editorial one. The model is driven by an audio input, so the language arrives with the file. As of 2026-09-22.
| What the documentation settles | What it leaves to a take |
|---|---|
| The card describes audio-driven video generation from an audio input with a reference image | Whether alignment holds equally across languages |
| Nothing is published about the language of the spoken performance | Whether any language is handled worse than others |
| A pose video argument lets the result follow a pose sequence while staying synchronised | Whether pose direction interacts with the sounds a language makes |
Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.
1A blank that the architecture accounts for
Most empty cells in this column are documents that did not get to the question. This one is a document for which the question barely arises: if a production supplies the waveform, the language is a property of the recording session rather than of the model. Publishing a list would be describing somebody else's inventory.
The register still records it as unstated, because the cell describes what the vendor published rather than what a reader can work out. But the reason behind the blank changes what a production should do about it, which is why the cell has a page.
2What a language could still affect here
Alignment does not treat every language alike. Mouth shapes differ, some languages carry more of their information in sounds a face barely moves for, and a model trained mostly on one language can align it better than another. None of that is addressed, and all of it is testable by anyone with the weights.
That is the compensation on this row. The licence is permissive and the weights are published, so a production that needs the answer can measure it rather than wait for a sentence.
3The other entries that never say which language
Eight more entries reach this column empty. This is the only one where the input itself supplies the answer, which is why the blank reads as a design rather than an omission.
- LTX Studio — not documented by the vendor.
- Luma Ray — not documented by the vendor.
- MiniMax — not documented by the vendor.
- Runway — not documented by the vendor.
- Sora 2 — not documented by the vendor.
- Veo — not documented by the vendor.
- Vidu — not documented by the vendor.
- Wan 3.0 — not documented by the vendor.
- LanguagesNot documented by the vendor (as of 2026-09-22)nothing published about the spoken performance
- Audio sourceThe card describes audio-driven cinematic video generation from an audio input with a reference image and an optional text prompthanded to the model rather than made by it
- A second control alongside the audioA pose video argument lets the result follow a pose sequence while staying synchronised to the audio
4Sources
Read from the model card at huggingface.co on 2026-09-22. The same column across every entry is on languages; everything this vendor publishes about speech is on Wan2.2-S2V. What counts as documented is on how read.