Sentioscope

Speech and voice controls, as each vendor documents them

Wan2.2-S2V: audio drives it, and the weights are out

The model card describes audio-driven cinematic video generation from an audio input with a reference image and an optional text prompt. It also states support for 480P and 720P and licenses the weights under Apache 2.0. As of 2026-09-22.

Three inputs arriving at one audio-driven generationThe card describes audio-driven cinematic video generation from an audio input with a reference image and an optional text prompt, and adds a pose video argument that lets the result follow a pose sequence.AudioThe driving inputSets the timingSynchronisation isthe promiseReference imageWho is on screenSets the subjectAppearance comesfrom herePose videoOptional sequenceSets the bodyA second channel ofdirectionText promptOptionalAdds descriptionLeast specified ofthe fourOne audio-driven generationWeights published under Apache 2.0, so blanks can be read from code
Fig. 1 The card does not say what happens when the pose track turns away from camera while the audio carries a line, and that conflict is easy to arrange.
Wan2.2-S2V on audio source, statement by statement. Read from the vendor's model card on 2026-09-22.
What the documentation settlesWhat it leaves to a take
Audio-driven cinematic video generation from an audio input with a reference imageHow well the motion reads when the audio carries two voices
Support is stated for 480P and 720P, with weights under Apache 2.0What behaviour a self-hosted deployment inherits that the card omits
A pose video argument lets the result follow a pose sequence while staying synchronisedHow a conflict between the pose track and the audio resolves

Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.

1Open weights change what a documented field can mean

Where a hosted API is the product, documentation is the contract and an absent sentence is a refusal to commit. Where the weights are published under a permissive licence, an absent sentence is often just a card that ran out of room, because anyone can read the code and find out.

This register still records the card rather than the code, for the same reason it does not record listening tests: the comparison only works if every entry is read the same way. The licence is logged as its own fact so a reader can weigh the blanks accordingly.

2A second control beside the audio, which nothing else here has

A pose argument means the body can be directed while the face stays tied to the waveform. That splits performance into two channels and is the closest thing in this register to conventional coverage, where blocking and delivery are decided separately.

The card does not say what happens when the two disagree, and disagreement is easy to arrange: a pose track that turns away from camera while the audio carries a line. Logged as a documented control with an undocumented precedence.

3The other entries that are handed a recording

Two more entries take the sound as an input. Neither of those publishes weights, so this is the only row here where the blanks can be resolved by reading code instead of a page.

  • Hedra — supplied; audio is required.
  • sync-3 — supplied, or read from text.
  • Audio source
    The card describes audio-driven cinematic video generation from an audio input with a reference image and an optional text prompthanded to the model rather than made by itWan-AI, Wan2.2-S2V-14B model card / recorded 2026-09-22
  • What the card does settle
    The card states support for 480P and 720P and licenses the weights under Apache 2.0Wan-AI, Wan2.2-S2V-14B model card / recorded 2026-09-22
  • A second control alongside the audio
    A pose video argument lets the result follow a pose sequence while staying synchronised to the audioWan-AI, Wan2.2-S2V-14B model card / recorded 2026-09-22

4Sources

Read from the model card at huggingface.co on 2026-09-22. The same column across every entry is on audio source; everything this vendor publishes about speech is on Wan2.2-S2V. What counts as documented is on how read.