Sora 2 and Luma Ray: a dialogue guide, or silence
Sora 2 has the most detailed dialogue guidance in this register: a place for the line, labels for speakers and an arithmetic for how much fits. Luma's video reference does not mention sound at all. As of 2026-09-22.
| Field | Sora 2 | Luma Ray | Where they part |
|---|---|---|---|
| Audio source | With the picture | Not documented by the vendor | Only Sora 2 answers |
| Voice source | Nothing documented as an input | Nothing documented as an input | Same answer |
| Per-character binding | On the prompt | Not documented by the vendor | Only Sora 2 answers |
| Languages | Not documented by the vendor | Not documented by the vendor | Same answer |
| Lip-sync | Named only where it fails | Not documented by the vendor | Only Sora 2 answers |
Inclusion rule. Two entries are given a page together when at least one column puts them at opposite grades of answer. Pairs that agree on every column, or that are both blank throughout, do not get a page. Order. Fixed field order, identical on every side-by-side page.
1The two extremes of the same column
One page tells a writer where dialogue belongs in a prompt, asks for speakers to be labelled consistently with turns alternated, and gives a rough word budget per shot length. The other names models and resolutions. Both are references for generating video.
Reading them together is the fastest way to see what this register measures. It is not which model sounds better; it is how much a production can decide before it starts spending.
2Detail in one column does not spread to the others
The detailed entry publishes no language, no voice source and no driver for the mouth. Its careful dialogue guidance never mentions which language a line is performed in, which is the first question a non-English production asks.
So the gap between these two narrows as the columns go on. On audio source they are at opposite ends; on voice source and languages they are both empty, for different reasons.
3A plan and a warning are both worth logging
The silent entry has one sentence on the subject anywhere: a page for AI assistants naming audio models as a planned integration. The detailed one has a warning that long, complex speeches are unlikely to sync.
Neither is a control. Both are the strongest thing their vendor has published in a column that is otherwise empty, which is exactly why this register records wording rather than ticks.
4Each of them on its own
The column this pair was chosen for is audio source, and each entry has a page of its own on it. The full row for either, all five columns with the wording behind each cell, is on its model note.
- Sora 2 on audio source — with the picture.
- Luma Ray on audio source — not documented by the vendor.
- Sora 2, all five fields — read from the model page and prompting guide.
- Luma Ray, all five fields — read from the video generation reference.
- How a spoken line is written into a promptDialogue belongs in a dialogue block below the prose description, with exchanges limited to a handful of sentences
- Dialogue timing against clip lengthA four-second shot is said to accommodate one or two short exchanges and an eight-second clip a few more, with long speeches unlikely to sync
- LanguagesNot documented by the vendor (as of 2026-09-22)neither a list nor a count
- Audio sourceNot documented by the vendor (as of 2026-09-22)no mention of sound in the video reference
- SoundPlans do integrate the following audio models into the Luma generative ecosystem: ElevenLabs SFX, Music and v3named as an integration to come
- What the video reference does settleThe video reference names ray-2 and ray-flash-2 and lists 540p, 720p, 1080 and 4k
5Sources
Every cell above is read from the documentation each vendor publishes, on the dates carried by the two model notes. The whole register on one column is on the field note; entries grouped by the shape of their answer are on the routes. Other pairs: Synthesia and D-ID, HeyGen and Synthesia.