Sora 2: synced audio, with the clock written down
The model page lists video and audio as output and describes generation with synced audio. The guide adds the clock: sixteen or twenty seconds a generation, with an extension adding up to twenty seconds at a time to a total of a hundred and twenty. As of 2026-09-22.
| What the documentation settles | What it leaves to a take |
|---|---|
| Videos are generated with synced audio, with both listed as output | How reliably the sync holds once a line runs long |
| Generations run sixteen or twenty seconds, extendable to a hundred and twenty in total | Whether an extension carries the same voice as the segment before it |
| A four-second shot holds one or two short exchanges; long speeches are unlikely to sync | Where exactly the alignment stops holding, since only a warning is given |
Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.
1A published length turns dialogue into arithmetic
Most entries in this register say audio comes out and stop. This one gives a writer both halves of a sum: how long a generation lasts, and roughly how much speech fits inside one. That is enough to size a scene on paper before a call is spent, which is the cheapest place to discover that a line does not fit.
The extension changes the shape of the arithmetic rather than removing it. Twenty seconds at a time to a total of two minutes is a sequence of joins, and every join is a place where a voice can drift, so the length is best read as a budget for a beat rather than for a scene.
2The warning is the closest thing here to a statement about mouths
Saying that long, complex speeches are unlikely to sync is an admission rather than a mechanism, and it is more useful than most mechanisms. It tells a writer to break a speech into exchanges, which is also what the prompting guide asks for when it wants turns alternated.
Read strictly, it is not enough to record lip-sync as documented, so that field stays marked unstated. Read practically, it is the sentence a shot list should be written against.
3The others that make sound while they make the picture
Seven more entries generate audio in the same pass. This is the only one that publishes both a clip length and guidance on how much speech to put in one.
- Kling AI — with the picture, on video 3.0.
- LTX Studio — with the picture, and audio to video.
- MiniMax — with the picture.
- SceneMixer — with the picture, in the language set for the project.
- Veo — with the picture.
- Vidu — with the picture, speech available on its own.
- Wan 3.0 — with the picture, unless switched off.
- Audio sourceDescribed as a media generation model generating videos with synced audio, with video and audio both listed as outputgenerated with the picture
- Shot length a line has to fitGenerations run 16 or 20 seconds, and an extension adds up to 20 seconds at a time to a total of 120
- Dialogue timing against clip lengthA four-second shot is said to accommodate one or two short exchanges and an eight-second clip a few more, with long speeches unlikely to sync
4Sources
Read from the model page and prompting guide at developers.openai.com on 2026-09-22. The same column across every entry is on audio source; everything this vendor publishes about speech is on Sora 2. What counts as documented is on how read.