Sora 2: lip-sync named only where it stops working
The only statement about alignment here is a warning: long, complex speeches are unlikely to sync. That is an admission about where the behaviour stops holding, and admissions of that kind are rarer than mechanisms. As of 2026-09-22.
| What the documentation settles | What it leaves to a take |
|---|---|
| Long, complex speeches are said to be unlikely to sync | Where exactly the threshold is, since only a direction is given |
| A four-second shot holds one or two short exchanges, an eight-second clip a few more | Whether short exchanges sync reliably rather than merely better |
| Videos are generated with synced audio, with both listed as output | What produces the sync, since no mechanism appears |
Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.
1A warning is more actionable than most claims in this column
Being told where something fails is often worth more than being told that it works. This sentence tells a writer to break a speech into exchanges, which is also what the guide separately asks for when it wants turns alternated, so two pieces of advice point the same way.
It is still not a mechanism. Nothing says what the model is aligning to, so a production cannot reason about which lines will be hardest; it can only follow the rule of thumb and keep the lines short.
2Why the field stays marked unstated anyway
Recording lip-sync as documented on the strength of a failure warning would overstate the page. The word appears, the limit is described, and the thing doing the work is not, so the cell keeps the grade that matches the evidence.
That is a judgement about reading rather than about the product. Sora 2 has the most detailed dialogue guidance in this register, and it is still the case that nobody reading it learns what moves a mouth.
3The others that name the feature and stop
Three more entries put lip-sync on the page without a driver. Those three name it as a capability; this one names it only to say when it will not hold.
- HeyGen — named, alongside translation.
- Kling AI — named, with nothing driving it.
- PixVerse — named as the endpoint's purpose.
- Dialogue timing against clip lengthA four-second shot is said to accommodate one or two short exchanges and an eight-second clip a few more, with long speeches unlikely to sync
- Shot length a line has to fitGenerations run 16 or 20 seconds, and an extension adds up to 20 seconds at a time to a total of 120
- Audio sourceDescribed as a media generation model generating videos with synced audio, with video and audio both listed as outputgenerated with the picture
4Sources
Read from the model page and prompting guide at developers.openai.com on 2026-09-22. The same column across every entry is on lip-sync; everything this vendor publishes about speech is on Sora 2. What counts as documented is on how read.