What Sora 2 documents about speech
Sora 2 (developers.openai.com) is described as generating videos with synced audio, and its prompting guide is the one place in this register that tells a writer where a line of dialogue goes and how to label two people talking. As of 2026-09-22.
| Field | What the vendor documents |
|---|---|
| Audio source | Generated with the picture; video and audio are both listed as output |
| Languages | Not documented |
| Voice source | Not documented |
| Per-character binding | Not documented; speakers are labelled inside the prompt |
| Lip-sync | Not named; long speeches are said to be unlikely to sync |
Inclusion rule. Fields are the same five for every model here; a field the vendor does not document is recorded as not documented rather than inferred from generated output. Order. Fixed field order, identical on every model page.
1Dialogue has a place in the prompt, and that is unusual
Most entries here describe audio as an output and stop there. This one tells a writer where the words go: spoken lines sit in a dialogue block underneath the prose description, so the model can tell a description of a room from a line said inside it. Exchanges are asked to stay within a handful of sentences.
For a shot list that is a format instruction rather than a quality claim. It means dialogue in a production's own documents can be written in the shape the model expects, and the conversion step that usually sits between a script and a prompt mostly disappears.
2Arithmetic a writer can plan against
The guide puts one or two short exchanges in a four-second shot and a few more in an eight-second one, and warns that long, complex speeches are unlikely to sync. Read beside the 16 and 20 second generation lengths, and the extension that adds up to 20 seconds at a time to a total of 120, that is enough to size a scene before anybody spends a call on it.
It is also as close as the page comes to a statement about mouths: not a mechanism, but an admission of where alignment stops holding. Recording lip-sync as documented on that basis would be over-reading it, so the field stays marked unstated.
3Whose voice it is stays unaddressed
Nothing on these pages says where a voice comes from, whether one can be supplied, or whether a character keeps one between calls. Speaker labels route lines to the right face inside a single generation; a label is an instruction within one call, not an identity that survives the next.
For a serial that is the gap that matters. A production can write a two-hander that plays correctly in one shot and still have no published way to make the same character sound the same in the following episode.
4This entry, one column at a time
Each of these stays inside a single field: what Sora 2 puts there, what the wording settles, and what it leaves for a take to answer.
- Audio source — with the picture.
- Per-character binding — on the prompt.
- Languages — not documented by the vendor.
- Lip-sync — named only where it fails.
5Read against another entry
Each of these puts Sora 2 beside one other entry on a column where the two land at opposite grades of answer.
- Sora 2 and Luma Ray — on audio source.
- Sora 2 and Wan 3.0 — on audio source.
- MiniMax and Sora 2 — on lip-sync.
6Where this model sits against the rest
| Model | Audio source |
|---|---|
| D-ID | A text script read aloud, or an audio url supplied |
| Hedra | Supplied; audio is a required input |
| HeyGen | A text to speech endpoint, apart from the picture |
| Kling AI | Generated with the picture, VIDEO 3.0 |
| LTX Studio | Generated jointly with the picture, plus audio-to-video |
| Luma Ray | Not documented; the video reference does not mention sound |
| MiniMax | Native speech, generated with the picture |
| PixVerse | Supplied to a speech endpoint, or read from text there |
| Runway | Not documented where voices are defined |
| SceneMixer | Generated with the picture, in the language set for the project |
| Sora 2 | Generated with the picture; video and audio are both listed as output |
| sync-3 | Supplied, or read from text on the call |
| Synthesia | Synthesised from the script, or an uploaded recording |
| Veo | Generated with the video rather than added afterwards |
| Vidu | Q3 generates speech, effects and music natively, on by default |
| Wan 3.0 | On by default; false returns a file with no audio track |
| Wan2.2-S2V | Supplied; the card describes the model as audio-driven |
Inclusion rule. Models whose vendor documents something about generated speech. A model with nothing published on the point is still listed, with the cell marked hollow. Order. Alphabetical by model name.
| Model | Languages |
|---|---|
| D-ID | A language field beside the voice; no list on the reference |
| Hedra | Full multi-language support claimed, none named |
| HeyGen | Thirty or more, counted on the translation endpoint |
| Kling AI | Five, not named |
| LTX Studio | Not documented |
| Luma Ray | Not documented |
| MiniMax | Not documented by the vendor |
| PixVerse | Multiple, with speech, singing and advertisements named as types |
| Runway | Not documented |
| SceneMixer | 15, named, with Cantonese a 16th for dialogue only |
| Sora 2 | Not documented |
| sync-3 | Ninety-five or more, counted rather than named |
| Synthesia | A catalogue, each voice carrying a language name and code |
| Veo | Not documented |
| Vidu | Not documented |
| Wan 3.0 | Not documented |
| Wan2.2-S2V | Not documented |
Inclusion rule. Models whose vendor documents something about generated speech. A model with nothing published on the point is still listed, with the cell marked hollow. Order. Alphabetical by model name.
| Model | Voice source |
|---|---|
| D-ID | A voice id from one of five named speech providers |
| Hedra | An uploaded track, or a voice id from the voices endpoint |
| HeyGen | A clone, instant from one recording or professional from 20 minutes or more |
| Kling AI | Not documented by the vendor |
| LTX Studio | Not documented |
| Luma Ray | Not documented |
| MiniMax | Reference audio, 15 s total across 3 clips |
| PixVerse | Built-in voices, or a custom voice from a supplied sample |
| Runway | An audio sample of 10 seconds to 5 minutes, or a description in words |
| SceneMixer | Preset voice library, or the user's own recordings |
| Sora 2 | Not documented |
| sync-3 | Not documented on the model page |
| Synthesia | The catalogue, or a cloned voice with a language list of its own |
| Veo | Not documented |
| Vidu | Not documented |
| Wan 3.0 | Reference audio, 15 seconds in total, WAV or MP3 up to 15 MB |
| Wan2.2-S2V | Whatever track is handed over; no catalogue |
Inclusion rule. Models whose vendor documents something about generated speech. A model with nothing published on the point is still listed, with the cell marked hollow. Order. Alphabetical by model name.
7Sources
Taken from the model page and prompting guide at developers.openai.com on 2026-09-22. Every entry against the same five fields sits on the speech controls page; the rule behind each field is on how read. The entries either side of this one: sync-3, Synthesia.
- Audio sourceDescribed as a media generation model generating videos with synced audio, with video and audio both listed as outputgenerated with the picture
- Shot length a line has to fitGenerations run 16 or 20 seconds, and an extension adds up to 20 seconds at a time to a total of 120
- How a spoken line is written into a promptDialogue belongs in a dialogue block below the prose description, with exchanges limited to a handful of sentences
- Two characters in frameFor multi-character scenes the guide asks for speakers labelled consistently and turns alternated, so each line lands on the right characteraddressed rather than skipped
- Dialogue timing against clip lengthA four-second shot is said to accommodate one or two short exchanges and an eight-second clip a few more, with long speeches unlikely to sync
- LanguagesNot documented by the vendor (as of 2026-09-22)neither a list nor a count