Entries that document a second lever beside the voice
Four entries document something nobody else here does. A pose sequence riding beside the audio. A mode returning speech alone. A route where a track drives the picture. A voice requested as a sentence rather than a sample. As of 2026-09-12.
| Model | What is documented | Detail |
|---|---|---|
| LTX Studio | Audio to video, the reverse direction | Beside joint generation |
| Runway | A voice described in at least twenty characters | Beside a sample route |
| Vidu | A speech-only mode | Beside three tracks on by default |
| Wan2.2-S2V | A pose video argument | Beside the driving audio |
Inclusion rule. Entries whose documentation describes a lever over speech or performance that no other entry in this register publishes. A control documented by several vendors belongs in a column instead. Order. Alphabetical by model name.
1A control documented once is a reason to read the page
Comparison tables are built from what vendors have in common, which makes them blind to the sentence that only one of them wrote. Each of these four is that sentence, and in each case it changes what the product is for rather than how well it performs.
None of them gets a column. A column of blanks with one entry filled is not a comparison, so these live on a page of their own and in the model notes they came from.
2Two of the four split a performance into channels
A pose track alongside audio separates blocking from delivery, which is how conventional coverage works and almost nothing generated manages. A reverse audio route makes the sound the fixed element and the picture the variable one, which inverts the usual order.
Both raise the same unanswered question: what happens when the two channels disagree. A pose that turns away while a line plays, or a track that implies action the prose does not, are easy to arrange and undocumented in both cases.
3The other two remove a constraint rather than adding a lever
A speech-only mode removes the generated bed from a deliverable. A described voice removes the need for a performer. Both make an otherwise ordinary product usable in a situation it would otherwise be shut out of.
That is worth as much as a capability. Most of the friction in generated dialogue is not a missing feature but a bundled one, and a documented way to switch part of a bundle off is rarer here than a new model line.
- Audio sourceJoint audio and video generation, plus audio-to-video, on LTX-2.5two directions
- Voice sourceA voice can instead be asked for in words, with the description required to run at least 20 charactersa route that needs no recording
- Audio sourceQ3 outputs speech, sound effects and music natively, on by default in the API, with a speech-only modethree tracks in one pass
- A second control alongside the audioA pose video argument lets the result follow a pose sequence while staying synchronised to the audio
4Sources
Every line is read from the vendor documentation linked on the speech controls page, checked 2026-09-12. What counts as documented is on how read. Nearby: What kind of page, Formats and file sizes, What sets the length.