What LTX Studio documents about speech
LTX Studio (ltx.io) documents joint audio and video generation on LTX-2.5 and an audio-to-video route in the other direction, and attaches voices to character Elements so a character keeps its voice between shots. As of 2026-09-12.
| Field | What the vendor documents |
|---|---|
| Audio source | Generated jointly with the picture, plus audio-to-video |
| Languages | Not documented |
| Voice source | Not documented |
| Per-character binding | Voices attached to character Elements |
| Lip-sync | Not documented |
Inclusion rule. Fields are the same five for every model here; a field the vendor does not document is recorded as not documented rather than inferred from generated output. Order. Fixed field order, identical on every model page.
1A documented route in both directions, which the others do not describe
Every other model here generates sound alongside the picture and stops there. This one also documents audio-to-video, where an existing track drives the generation rather than being produced by it. For advertising and music work that inverts the usual order of operations: the audio is the fixed element and the picture is fitted to it.
What it means for drama is less obvious and the documentation does not say. A dialogue track recorded by an actor is an existing track, and whether the route is intended to accept one is not addressed anywhere public. It is recorded here as a documented capability rather than as a workflow the vendor endorses.

2Voices live on the character, which is the durable form
Attaching a voice to a character Element puts it in the same place as the character's appearance, so a production defines the cast once and draws on it. That is the arrangement that survives a long shoot, because nothing has to be re-specified per shot and there is no per-call step to forget.
Three fields stay empty against it. No language list, no statement about where the voices themselves come from, and nothing on lip-sync. A production can therefore expect a character to sound like itself, without being able to say in advance what it will sound like or in which languages it can speak.
3This entry, one column at a time
Each of these stays inside a single field: what LTX Studio puts there, what the wording settles, and what it leaves for a take to answer.
- Audio source — with the picture, and audio to video.
- Voice source — nothing documented as an input.
- Per-character binding — on the character element.
4Read against another entry
Each of these puts LTX Studio beside one other entry on a column where the two land at opposite grades of answer.
- HeyGen and LTX Studio — on per-character binding.
- LTX Studio and Vidu — on per-character binding.
5Where this model sits against the rest
| Model | Audio source |
|---|---|
| D-ID | A text script read aloud, or an audio url supplied |
| Hedra | Supplied; audio is a required input |
| HeyGen | A text to speech endpoint, apart from the picture |
| Kling AI | Generated with the picture, VIDEO 3.0 |
| LTX Studio | Generated jointly with the picture, plus audio-to-video |
| Luma Ray | Not documented; the video reference does not mention sound |
| MiniMax | Native speech, generated with the picture |
| PixVerse | Supplied to a speech endpoint, or read from text there |
| Runway | Not documented where voices are defined |
| SceneMixer | Generated with the picture, in the language set for the project |
| Sora 2 | Generated with the picture; video and audio are both listed as output |
| sync-3 | Supplied, or read from text on the call |
| Synthesia | Synthesised from the script, or an uploaded recording |
| Veo | Generated with the video rather than added afterwards |
| Vidu | Q3 generates speech, effects and music natively, on by default |
| Wan 3.0 | On by default; false returns a file with no audio track |
| Wan2.2-S2V | Supplied; the card describes the model as audio-driven |
Inclusion rule. Models whose vendor documents something about generated speech. A model with nothing published on the point is still listed, with the cell marked hollow. Order. Alphabetical by model name.
| Model | Languages |
|---|---|
| D-ID | A language field beside the voice; no list on the reference |
| Hedra | Full multi-language support claimed, none named |
| HeyGen | Thirty or more, counted on the translation endpoint |
| Kling AI | Five, not named |
| LTX Studio | Not documented |
| Luma Ray | Not documented |
| MiniMax | Not documented by the vendor |
| PixVerse | Multiple, with speech, singing and advertisements named as types |
| Runway | Not documented |
| SceneMixer | 15, named, with Cantonese a 16th for dialogue only |
| Sora 2 | Not documented |
| sync-3 | Ninety-five or more, counted rather than named |
| Synthesia | A catalogue, each voice carrying a language name and code |
| Veo | Not documented |
| Vidu | Not documented |
| Wan 3.0 | Not documented |
| Wan2.2-S2V | Not documented |
Inclusion rule. Models whose vendor documents something about generated speech. A model with nothing published on the point is still listed, with the cell marked hollow. Order. Alphabetical by model name.
| Model | Voice source |
|---|---|
| D-ID | A voice id from one of five named speech providers |
| Hedra | An uploaded track, or a voice id from the voices endpoint |
| HeyGen | A clone, instant from one recording or professional from 20 minutes or more |
| Kling AI | Not documented by the vendor |
| LTX Studio | Not documented |
| Luma Ray | Not documented |
| MiniMax | Reference audio, 15 s total across 3 clips |
| PixVerse | Built-in voices, or a custom voice from a supplied sample |
| Runway | An audio sample of 10 seconds to 5 minutes, or a description in words |
| SceneMixer | Preset voice library, or the user's own recordings |
| Sora 2 | Not documented |
| sync-3 | Not documented on the model page |
| Synthesia | The catalogue, or a cloned voice with a language list of its own |
| Veo | Not documented |
| Vidu | Not documented |
| Wan 3.0 | Reference audio, 15 seconds in total, WAV or MP3 up to 15 MB |
| Wan2.2-S2V | Whatever track is handed over; no catalogue |
Inclusion rule. Models whose vendor documents something about generated speech. A model with nothing published on the point is still listed, with the cell marked hollow. Order. Alphabetical by model name.
6Sources
Taken from the product page at ltx.io on 2026-09-12. Every entry against the same five fields sits on the speech controls page; the rule behind each field is on how read. The entries either side of this one: Luma Ray, MiniMax.
- Audio sourceJoint audio and video generation, plus audio-to-video, on LTX-2.5two directions
- Per-character bindingVoices are attached to character Elementstied to the element system