Sentioscope

Speech and voice controls, as each vendor documents them

LTX Studio: sound and picture, in both directions

LTX Studio documents two directions at once. Audio and video generated jointly on LTX-2.5, and audio-to-video, where an existing track drives the generation rather than coming out of it. Nothing else in this register describes both. As of 2026-09-12.

Sound out of the model, or sound into itJoint audio and video generation on LTX-2.5 makes the sound an output. The audio-to-video route makes it an input. The two arrangements suit opposite kinds of work and are documented on the same page.Joint generationAudio to videoWhat is fixed firstThe prose description of the shotThe finished trackWhat the model decidesThe performance that comes backThe picture cut to that trackWhere it fitsDrama, where the words are writtenAdvertising and music, where the audio issetWhat the page settlesBoth run on the same productBoth run on the same productOne page, two orders of operation
Fig. 1 The page does not say whether a recorded dialogue track is what the second route expects, so the intent is logged as the vendor left it.
LTX Studio on audio source, statement by statement. Read from the vendor's product page on 2026-09-12.
What the documentation settlesWhat it leaves to a take
Joint audio and video generation is documented on LTX-2.5Whether the joint route can be steered towards a particular voice
An audio-to-video route runs in the other directionWhether a recorded dialogue track is what that route expects
Voices are attached to character ElementsWhere those voices come from, since no source is published

Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.

1Two directions are two different products sharing a page

Generating sound with the picture and fitting a picture to sound are opposite arrangements. The first suits drama, where the words are written and the performance is whatever the model gives back. The second suits advertising and music, where the audio is the fixed element and the picture is cut to it.

Having both documented in one place is unusual and mildly ambiguous, because the page does not say whether a dialogue recording counts as the kind of audio the second route wants. This register logs the capability and leaves the intent where the vendor left it.

2The gap is everything about the voice itself

Voices attach to character Elements, which is the durable arrangement, and that is where the page stops. No language list, no statement about what a voice can be built from, and nothing about mouths. A production can expect a character to sound like itself without being able to say in advance what that is.

Read together with the audio-to-video route, the absence has a plausible shape: a platform that lets a track drive the picture has less reason to document voice construction, because a team with strong opinions about the voice can simply supply one.

3The others that make sound while they make the picture

Seven more entries generate audio in the same pass as the shot. This is the only one of them that also documents the reverse direction.

  • Kling AI — with the picture, on video 3.0.
  • MiniMax — with the picture.
  • SceneMixer — with the picture, in the language set for the project.
  • Sora 2 — with the picture.
  • Veo — with the picture.
  • Vidu — with the picture, speech available on its own.
  • Wan 3.0 — with the picture, unless switched off.
  • Audio source
    Joint audio and video generation, plus audio-to-video, on LTX-2.5two directionsLTX Studio / recorded 2026-09-12
  • Per-character binding
    Voices are attached to character Elementstied to the element systemLTX Studio / recorded 2026-09-12

4Sources

Read from the product page at ltx.io on 2026-09-12. The same column across every entry is on audio source; everything this vendor publishes about speech is on LTX Studio. What counts as documented is on how read.