What Luma Ray documents about speech
Luma Ray (docs.lumalabs.ai) documents ray-2 and ray-flash-2 without mentioning sound anywhere in the video reference. The nearest audio statement found is a note that ElevenLabs sound effects, music and a voice model are planned integrations. As of 2026-09-22.
| Field | What the vendor documents |
|---|---|
| Audio source | Not documented; the video reference does not mention sound |
| Languages | Not documented |
| Voice source | Not documented |
| Per-character binding | Not documented |
| Lip-sync | Not documented |
Inclusion rule. Fields are the same five for every model here; a field the vendor does not document is recorded as not documented rather than inferred from generated output. Order. Fixed field order, identical on every model page.
1An entry with every field empty, kept for exactly that reason
The video reference is precise about what it covers: two model names, resolutions from 540p to 4k, keyframes, a duration in the request body. Sound is not among them. Read end to end it describes a picture generator, and a production planning dialogue leaves the page with nothing to plan.
Dropping the row would be the easy choice and the wrong one. A register of documented speech controls that listed only vendors with speech controls would report a market in which every model talks, and that is not the market these notes found.
2A plan is not a control
The one official sentence about audio sits on a page written for answer engines, and it looks forward: sound effects, music and a voice model are named as integrations planned into the ecosystem. That is a roadmap, and no field here is filled from one.
The distinction is not pedantry. An integration arriving next quarter helps a series that starts next year and does nothing for one shooting now, so a table recording the two alike would mislead in the direction that costs money.
3What the silence most likely means
Sound named as a partner integration rather than as a model parameter points to audio being assembled after the picture, in an edit. That is a defensible architecture and a different production shape from the entries here that return a finished mix, because a scoring or dubbing stage has to be scheduled and paid for.
None of that is published, so none of it is recorded as fact. What is recorded is that on the day these pages were read, a buyer could not learn from the vendor whether a generated shot would contain a voice at all.
4This entry, one column at a time
Each of these stays inside a single field: what Luma Ray puts there, what the wording settles, and what it leaves for a take to answer.
- Audio source — not documented by the vendor.
5Read against another entry
Each of these puts Luma Ray beside one other entry on a column where the two land at opposite grades of answer.
- Sora 2 and Luma Ray — on audio source.
- Runway and Luma Ray — on voice source.
6Where this model sits against the rest
| Model | Audio source |
|---|---|
| D-ID | A text script read aloud, or an audio url supplied |
| Hedra | Supplied; audio is a required input |
| HeyGen | A text to speech endpoint, apart from the picture |
| Kling AI | Generated with the picture, VIDEO 3.0 |
| LTX Studio | Generated jointly with the picture, plus audio-to-video |
| Luma Ray | Not documented; the video reference does not mention sound |
| MiniMax | Native speech, generated with the picture |
| PixVerse | Supplied to a speech endpoint, or read from text there |
| Runway | Not documented where voices are defined |
| SceneMixer | Generated with the picture, in the language set for the project |
| Sora 2 | Generated with the picture; video and audio are both listed as output |
| sync-3 | Supplied, or read from text on the call |
| Synthesia | Synthesised from the script, or an uploaded recording |
| Veo | Generated with the video rather than added afterwards |
| Vidu | Q3 generates speech, effects and music natively, on by default |
| Wan 3.0 | On by default; false returns a file with no audio track |
| Wan2.2-S2V | Supplied; the card describes the model as audio-driven |
Inclusion rule. Models whose vendor documents something about generated speech. A model with nothing published on the point is still listed, with the cell marked hollow. Order. Alphabetical by model name.
| Model | Languages |
|---|---|
| D-ID | A language field beside the voice; no list on the reference |
| Hedra | Full multi-language support claimed, none named |
| HeyGen | Thirty or more, counted on the translation endpoint |
| Kling AI | Five, not named |
| LTX Studio | Not documented |
| Luma Ray | Not documented |
| MiniMax | Not documented by the vendor |
| PixVerse | Multiple, with speech, singing and advertisements named as types |
| Runway | Not documented |
| SceneMixer | 15, named, with Cantonese a 16th for dialogue only |
| Sora 2 | Not documented |
| sync-3 | Ninety-five or more, counted rather than named |
| Synthesia | A catalogue, each voice carrying a language name and code |
| Veo | Not documented |
| Vidu | Not documented |
| Wan 3.0 | Not documented |
| Wan2.2-S2V | Not documented |
Inclusion rule. Models whose vendor documents something about generated speech. A model with nothing published on the point is still listed, with the cell marked hollow. Order. Alphabetical by model name.
| Model | Voice source |
|---|---|
| D-ID | A voice id from one of five named speech providers |
| Hedra | An uploaded track, or a voice id from the voices endpoint |
| HeyGen | A clone, instant from one recording or professional from 20 minutes or more |
| Kling AI | Not documented by the vendor |
| LTX Studio | Not documented |
| Luma Ray | Not documented |
| MiniMax | Reference audio, 15 s total across 3 clips |
| PixVerse | Built-in voices, or a custom voice from a supplied sample |
| Runway | An audio sample of 10 seconds to 5 minutes, or a description in words |
| SceneMixer | Preset voice library, or the user's own recordings |
| Sora 2 | Not documented |
| sync-3 | Not documented on the model page |
| Synthesia | The catalogue, or a cloned voice with a language list of its own |
| Veo | Not documented |
| Vidu | Not documented |
| Wan 3.0 | Reference audio, 15 seconds in total, WAV or MP3 up to 15 MB |
| Wan2.2-S2V | Whatever track is handed over; no catalogue |
Inclusion rule. Models whose vendor documents something about generated speech. A model with nothing published on the point is still listed, with the cell marked hollow. Order. Alphabetical by model name.
7Sources
Taken from the video generation reference at docs.lumalabs.ai on 2026-09-22. Every entry against the same five fields sits on the speech controls page; the rule behind each field is on how read. The entries either side of this one: MiniMax, PixVerse.
- Audio sourceNot documented by the vendor (as of 2026-09-22)no mention of sound in the video reference
- SoundPlans do integrate the following audio models into the Luma generative ecosystem: ElevenLabs SFX, Music and v3named as an integration to come
- What the video reference does settleThe video reference names ray-2 and ray-flash-2 and lists 540p, 720p, 1080 and 4k