LTX Studio: sound and picture, in both directions
LTX Studio documents two directions at once. Audio and video generated jointly on LTX-2.5, and audio-to-video, where an existing track drives the generation rather than coming out of it. Nothing else in this register describes both. As of 2026-09-12.
| What the documentation settles | What it leaves to a take |
|---|---|
| Joint audio and video generation is documented on LTX-2.5 | Whether the joint route can be steered towards a particular voice |
| An audio-to-video route runs in the other direction | Whether a recorded dialogue track is what that route expects |
| Voices are attached to character Elements | Where those voices come from, since no source is published |
Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.
1Two directions are two different products sharing a page
Generating sound with the picture and fitting a picture to sound are opposite arrangements. The first suits drama, where the words are written and the performance is whatever the model gives back. The second suits advertising and music, where the audio is the fixed element and the picture is cut to it.
Having both documented in one place is unusual and mildly ambiguous, because the page does not say whether a dialogue recording counts as the kind of audio the second route wants. This register logs the capability and leaves the intent where the vendor left it.
2The gap is everything about the voice itself
Voices attach to character Elements, which is the durable arrangement, and that is where the page stops. No language list, no statement about what a voice can be built from, and nothing about mouths. A production can expect a character to sound like itself without being able to say in advance what that is.
Read together with the audio-to-video route, the absence has a plausible shape: a platform that lets a track drive the picture has less reason to document voice construction, because a team with strong opinions about the voice can simply supply one.
3The others that make sound while they make the picture
Seven more entries generate audio in the same pass as the shot. This is the only one of them that also documents the reverse direction.
- Kling AI — with the picture, on video 3.0.
- MiniMax — with the picture.
- SceneMixer — with the picture, in the language set for the project.
- Sora 2 — with the picture.
- Veo — with the picture.
- Vidu — with the picture, speech available on its own.
- Wan 3.0 — with the picture, unless switched off.
- Audio sourceJoint audio and video generation, plus audio-to-video, on LTX-2.5two directions
- Per-character bindingVoices are attached to character Elementstied to the element system
4Sources
Read from the product page at ltx.io on 2026-09-12. The same column across every entry is on audio source; everything this vendor publishes about speech is on LTX Studio. What counts as documented is on how read.