What Wan2.2-S2V documents about speech
Wan2.2-S2V (huggingface.co) is an audio-driven model with open weights: a track, a reference image and an optional prompt go in, 480P or 720P comes out, and the weights carry an Apache 2.0 licence. As of 2026-09-22.
| Field | What the vendor documents |
|---|---|
| Audio source | Supplied; the card describes the model as audio-driven |
| Languages | Not documented |
| Voice source | Whatever track is handed over; no catalogue |
| Per-character binding | Not documented |
| Lip-sync | The result is synchronised to the audio input |
Inclusion rule. Fields are the same five for every model here; a field the vendor does not document is recorded as not documented rather than inferred from generated output. Order. Fixed field order, identical on every model page.
1Open weights change what a documented field can mean
This is the one entry here whose documentation is a model card rather than a product page, and the difference shows in what gets settled. Resolutions, the licence, the arguments and the inputs are all stated; voices, languages and a cast are not, because a card describes a model and not a service wrapped around one.
Apache 2.0 also moves the compliance question. A production running the weights itself decides what audio goes in, which is a stronger position than a terms page and a heavier obligation, since nothing upstream is checking the consent behind a voice.
2A pose track beside the audio
A pose video argument lets the result follow a pose sequence while staying synchronised to the audio. Two drivers at once is a more filmic arrangement than audio alone: the body can be directed while the mouth follows the take, which is roughly how a performance-capture pipeline is organised.
For a serial that matters more than it sounds. Most audio-driven entries here animate a portrait; a pose track is what makes a shot rather than a bust, and it is published as an argument rather than as a promise.
3Two fields empty by construction
Languages and voice source are blank here, and not because the card is thin. The audio is supplied, so the language is whatever was recorded and the voice is whoever recorded it, neither of which the model has reason to describe. The register still marks them unstated, because a reader comparing rows needs to know the answer is not on the page.
What is left is a clean division of labour. Casting, language and consent are settled upstream in a recording session, and the model is asked only to make a face agree with a waveform.
4This entry, one column at a time
Each of these stays inside a single field: what Wan2.2-S2V puts there, what the wording settles, and what it leaves for a take to answer.
- Audio source — supplied; the model is audio-driven.
- Voice source — the audio track itself.
- Languages — not documented by the vendor.
- Lip-sync — the audio input.
5Read against another entry
Each of these puts Wan2.2-S2V beside one other entry on a column where the two land at opposite grades of answer.
- Wan 3.0 and Wan2.2-S2V — on audio source.
6Where this model sits against the rest
| Model | Audio source |
|---|---|
| D-ID | A text script read aloud, or an audio url supplied |
| Hedra | Supplied; audio is a required input |
| HeyGen | A text to speech endpoint, apart from the picture |
| Kling AI | Generated with the picture, VIDEO 3.0 |
| LTX Studio | Generated jointly with the picture, plus audio-to-video |
| Luma Ray | Not documented; the video reference does not mention sound |
| MiniMax | Native speech, generated with the picture |
| PixVerse | Supplied to a speech endpoint, or read from text there |
| Runway | Not documented where voices are defined |
| SceneMixer | Generated with the picture, in the language set for the project |
| Sora 2 | Generated with the picture; video and audio are both listed as output |
| sync-3 | Supplied, or read from text on the call |
| Synthesia | Synthesised from the script, or an uploaded recording |
| Veo | Generated with the video rather than added afterwards |
| Vidu | Q3 generates speech, effects and music natively, on by default |
| Wan 3.0 | On by default; false returns a file with no audio track |
| Wan2.2-S2V | Supplied; the card describes the model as audio-driven |
Inclusion rule. Models whose vendor documents something about generated speech. A model with nothing published on the point is still listed, with the cell marked hollow. Order. Alphabetical by model name.
| Model | Languages |
|---|---|
| D-ID | A language field beside the voice; no list on the reference |
| Hedra | Full multi-language support claimed, none named |
| HeyGen | Thirty or more, counted on the translation endpoint |
| Kling AI | Five, not named |
| LTX Studio | Not documented |
| Luma Ray | Not documented |
| MiniMax | Not documented by the vendor |
| PixVerse | Multiple, with speech, singing and advertisements named as types |
| Runway | Not documented |
| SceneMixer | 15, named, with Cantonese a 16th for dialogue only |
| Sora 2 | Not documented |
| sync-3 | Ninety-five or more, counted rather than named |
| Synthesia | A catalogue, each voice carrying a language name and code |
| Veo | Not documented |
| Vidu | Not documented |
| Wan 3.0 | Not documented |
| Wan2.2-S2V | Not documented |
Inclusion rule. Models whose vendor documents something about generated speech. A model with nothing published on the point is still listed, with the cell marked hollow. Order. Alphabetical by model name.
| Model | Voice source |
|---|---|
| D-ID | A voice id from one of five named speech providers |
| Hedra | An uploaded track, or a voice id from the voices endpoint |
| HeyGen | A clone, instant from one recording or professional from 20 minutes or more |
| Kling AI | Not documented by the vendor |
| LTX Studio | Not documented |
| Luma Ray | Not documented |
| MiniMax | Reference audio, 15 s total across 3 clips |
| PixVerse | Built-in voices, or a custom voice from a supplied sample |
| Runway | An audio sample of 10 seconds to 5 minutes, or a description in words |
| SceneMixer | Preset voice library, or the user's own recordings |
| Sora 2 | Not documented |
| sync-3 | Not documented on the model page |
| Synthesia | The catalogue, or a cloned voice with a language list of its own |
| Veo | Not documented |
| Vidu | Not documented |
| Wan 3.0 | Reference audio, 15 seconds in total, WAV or MP3 up to 15 MB |
| Wan2.2-S2V | Whatever track is handed over; no catalogue |
Inclusion rule. Models whose vendor documents something about generated speech. A model with nothing published on the point is still listed, with the cell marked hollow. Order. Alphabetical by model name.
7Sources
Taken from the model card at huggingface.co on 2026-09-22. Every entry against the same five fields sits on the speech controls page; the rule behind each field is on how read. The entries either side of this one: D-ID, Hedra.
- Audio sourceThe card describes audio-driven cinematic video generation from an audio input with a reference image and an optional text prompthanded to the model rather than made by it
- What the card does settleThe card states support for 480P and 720P and licenses the weights under Apache 2.0
- A second control alongside the audioA pose video argument lets the result follow a pose sequence while staying synchronised to the audio
- LanguagesNot documented by the vendor (as of 2026-09-22)nothing published about the spoken performance