LTX Studio and Vidu: a bound voice, or none described
Two native-audio entries, one of which has thought about a cast. Voices attach to character Elements on one; the other documents speech, effects and music as outputs and says nothing about whose voice it is. As of 2026-09-22.
| Field | LTX Studio | Vidu | Where they part |
|---|---|---|---|
| Audio source | With the picture, and audio to video | With the picture, speech available on its own | Different answers |
| Voice source | Nothing documented as an input | Nothing documented as an input | Same answer |
| Per-character binding | On the character Element | Not documented by the vendor | Only LTX Studio answers |
| Languages | Not documented by the vendor | Not documented by the vendor | Same answer |
| Lip-sync | Not documented by the vendor | Not documented by the vendor | Same answer |
Inclusion rule. Two entries are given a page together when at least one column puts them at opposite grades of answer. Pairs that agree on every column, or that are both blank throughout, do not get a page. Order. Fixed field order, identical on every side-by-side page.
1Describing what is produced, against describing who speaks
One reference tells a reader what comes out of a call: three tracks, on by default, with a mode that isolates speech. The other tells a reader where a voice lives: on the character, travelling between generations.
For one clip the first is enough. For a series the second is the question, because a cast re-drawn on every call is not a cast, and no amount of detail about output tracks substitutes for it.
2Both leave the source of the voice unstated
Neither publishes anything a production could hand over. One binds a voice to a character without saying where the voice came from; the other does not mention voices at all. So on selection they are equally silent and on persistence they are at opposite ends.
That split is worth seeing, because a feature table with one column for voices would score these two the same. Five columns exist so that persistence and selection cannot be averaged into a tick.
3One documents a reverse route, the other a clean stem
Beyond this register's five columns each has one genuinely distinctive sentence. One documents audio-to-video, where an existing track drives the generation. The other documents a speech-only mode that keeps dialogue and drops the bed.
Both are the kind of affordance a drama pipeline actually uses, and neither vendor publishes the other's. A production wanting both would be choosing which stage to solve outside the tool.
4Each of them on its own
The column this pair was chosen for is per-character binding, and each entry has a page of its own on it. The full row for either, all five columns with the wording behind each cell, is on its model note.
- LTX Studio on per-character binding — on the character element.
- Vidu on per-character binding — not documented by the vendor.
- LTX Studio, all five fields — read from the product page.
- Vidu, all five fields — read from the API reference.
- Per-character bindingVoices are attached to character Elementstied to the element system
- Audio sourceJoint audio and video generation, plus audio-to-video, on LTX-2.5two directions
- Audio sourceQ3 outputs speech, sound effects and music natively, on by default in the API, with a speech-only modethree tracks in one pass
- Per-character bindingNot documented by the vendor (as of 2026-09-22)no public description
5Sources
Every cell above is read from the documentation each vendor publishes, on the dates carried by the two model notes. The whole register on one column is on the field note; entries grouped by the shape of their answer are on the routes. Other pairs: MiniMax and Sora 2, Hedra and Synthesia.