Sentioscope

Speech and voice controls, as each vendor documents them

LTX Studio and Vidu: a bound voice, or none described

Two native-audio entries, one of which has thought about a cast. Voices attach to character Elements on one; the other documents speech, effects and music as outputs and says nothing about whose voice it is. As of 2026-09-22.

Describing what comes out, or describing who speaksOne reference says what a call returns: three tracks, on by default, with a mode that isolates speech. The other says where a voice lives: on the character, travelling between generations.LTX StudioViduA voice lastingAttached to a character ElementNot documentedWhere a voice comes fromNot documentedNot documentedDistinctive sentenceAudio to video, the reverse directionA speech-only modeFor one clipEnoughEnoughOpposite ends on persistence, equally silent on selection
Fig. 1 A feature table with one column for voices would score these two the same, which is why five columns exist.
Controls documented, model by modelFilled where the model documents that control, hollow where nothing is published about it.Controls documented, model by modelLTX StudioViduWhere they partAudio sourceAudio source — LTX Studio: With the picture, and audio to videoAudio source — Vidu: With the picture, speech available on its ownAudio source — Where they part: Different answersVoice sourceVoice source — LTX Studio: Nothing documented as an inputVoice source — Vidu: Nothing documented as an inputVoice source — Where they part: Same answerPer-character bindingPer-character binding — LTX Studio: On the character ElementPer-character binding — Vidu: Not documented by the vendorPer-character binding — Where they part: Only LTX Studio answersLanguagesLanguages — LTX Studio: Not documented by the vendorLanguages — Vidu: Not documented by the vendorLanguages — Where they part: Same answerLip-syncLip-sync — LTX Studio: Not documented by the vendorLip-sync — Vidu: Not documented by the vendorLip-sync — Where they part: Same answer
Fig. 2 Filled where the model documents that control, hollow where nothing is published about it.
Models documenting each controlHow many models document each control. A hollow column is a statement about documentation, not capability.Models documenting each controlLTX Studio3 of 5Vidu2 of 5Where they part5 of 5
Fig. 3 How many models document each control. A hollow column is a statement about documentation, not capability.
LTX Studio and Vidu on the five fields this register keeps, with the gap on each. Recorded 2026-09-22.
FieldLTX StudioViduWhere they part
Audio sourceWith the picture, and audio to videoWith the picture, speech available on its ownDifferent answers
Voice sourceNothing documented as an inputNothing documented as an inputSame answer
Per-character bindingOn the character ElementNot documented by the vendorOnly LTX Studio answers
LanguagesNot documented by the vendorNot documented by the vendorSame answer
Lip-syncNot documented by the vendorNot documented by the vendorSame answer

Inclusion rule. Two entries are given a page together when at least one column puts them at opposite grades of answer. Pairs that agree on every column, or that are both blank throughout, do not get a page. Order. Fixed field order, identical on every side-by-side page.

1Describing what is produced, against describing who speaks

One reference tells a reader what comes out of a call: three tracks, on by default, with a mode that isolates speech. The other tells a reader where a voice lives: on the character, travelling between generations.

For one clip the first is enough. For a series the second is the question, because a cast re-drawn on every call is not a cast, and no amount of detail about output tracks substitutes for it.

2Both leave the source of the voice unstated

Neither publishes anything a production could hand over. One binds a voice to a character without saying where the voice came from; the other does not mention voices at all. So on selection they are equally silent and on persistence they are at opposite ends.

That split is worth seeing, because a feature table with one column for voices would score these two the same. Five columns exist so that persistence and selection cannot be averaged into a tick.

3One documents a reverse route, the other a clean stem

Beyond this register's five columns each has one genuinely distinctive sentence. One documents audio-to-video, where an existing track drives the generation. The other documents a speech-only mode that keeps dialogue and drops the bed.

Both are the kind of affordance a drama pipeline actually uses, and neither vendor publishes the other's. A production wanting both would be choosing which stage to solve outside the tool.

4Each of them on its own

The column this pair was chosen for is per-character binding, and each entry has a page of its own on it. The full row for either, all five columns with the wording behind each cell, is on its model note.

  • Per-character binding
    Voices are attached to character Elementstied to the element systemLTX Studio / recorded 2026-09-12
  • Audio source
    Joint audio and video generation, plus audio-to-video, on LTX-2.5two directionsLTX Studio / recorded 2026-09-12
  • Audio source
    Q3 outputs speech, sound effects and music natively, on by default in the API, with a speech-only modethree tracks in one passVidu, API reference / recorded 2026-09-12
  • Per-character binding
    Not documented by the vendor (as of 2026-09-22)no public descriptionVidu, API reference / recorded 2026-09-12

5Sources

Every cell above is read from the documentation each vendor publishes, on the dates carried by the two model notes. The whole register on one column is on the field note; entries grouped by the shape of their answer are on the routes. Other pairs: MiniMax and Sora 2, Hedra and Synthesia.