Sentioscope

Speech and voice controls, as each vendor documents them

Wan2.2-S2V: the voice is the file you hand over

There is no catalogue and no sample window, because the audio is not a reference for a voice but the performance itself. The card describes audio-driven generation from an audio input with a reference image and an optional prompt. As of 2026-09-22.

The clip is the delivery, not an example to copyEvery other entry taking audio here treats the clip as something for the model to imitate. On this card the audio is the performance the audience will hear, and the model's job is the picture around it.A reference clipA supplied performanceWhat the model does with itImitates the voice in itKeeps it and builds a picturePublished ceilingSeconds, and often megabytesNone to publishWhat can go wrongThe copy misses accent or pacingNothing about the voice; only the pictureWhere casting happensIn the choice of excerptIn the recording sessionWeights under Apache 2.0, so the blanks can be read from code
Fig. 1 Casting moves entirely into a recording session, where the tools are microphones and performers rather than parameters.
Wan2.2-S2V on voice source, statement by statement. Read from the vendor's model card on 2026-09-22.
What the documentation settlesWhat it leaves to a take
The card describes audio-driven video generation from an audio input with a reference imageHow the result behaves when the recording carries two voices
Support is stated for 480P and 720P, with weights published under Apache 2.0What a self-hosted deployment inherits that the card omits
A pose video argument lets the result follow a pose sequence while staying synchronisedHow a conflict between pose and audio resolves

Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.

1A supplied performance is not a supplied reference

Every other entry taking audio in this column treats the clip as an example for the model to imitate. Here the clip is the delivery: whatever was recorded is what the audience hears, and the model's job is the picture. That removes the whole class of questions about how faithfully a voice is copied.

It also moves the casting entirely into a recording session, where the tools are microphones and performers rather than parameters. A production that already works this way loses nothing from an empty catalogue.

2Permissive weights change what an empty cell is worth

On a hosted API, an absent sentence is a refusal to commit. Here the weights are published under Apache 2.0, so a team that needs to know how the model treats a particular kind of recording can find out by running it rather than by asking.

The register still records the card, because a comparison only works if every entry is read the same way. The licence is logged as its own fact so a reader can discount this row's blanks accordingly.

3The others that take a clip on the call that uses it

Two more entries take audio with each generation. Both treat it as a reference with a fifteen-second ceiling; this one takes the finished performance and has no ceiling to publish.

  • MiniMax — a reference clip on the call.
  • Wan 3.0 — reference audio, 15 seconds in total.
  • Audio source
    The card describes audio-driven cinematic video generation from an audio input with a reference image and an optional text prompthanded to the model rather than made by itWan-AI, Wan2.2-S2V-14B model card / recorded 2026-09-22
  • What the card does settle
    The card states support for 480P and 720P and licenses the weights under Apache 2.0Wan-AI, Wan2.2-S2V-14B model card / recorded 2026-09-22
  • A second control alongside the audio
    A pose video argument lets the result follow a pose sequence while staying synchronised to the audioWan-AI, Wan2.2-S2V-14B model card / recorded 2026-09-22

4Sources

Read from the model card at huggingface.co on 2026-09-22. The same column across every entry is on voice source; everything this vendor publishes about speech is on Wan2.2-S2V. What counts as documented is on how read.