What happens with two speakers in frame
One vendor now addresses it: the Sora 2 prompting guide asks for speakers to be labelled consistently and turns alternated, so each line lands on the right character. Two others say something adjacent. Everybody else is silent on the shot a dialogue scene is built from. As of 2026-09-12.
| Model | What is documented | Detail |
|---|---|---|
| Documented guidance | Label speakers, alternate turns | Sora 2 prompting guide |
| Nearest mechanism | Lip-sync tied to the on-screen speaker | MiniMax video guide |
| A framing condition | Lip sync performs better framed closer | Synthesia avatar docs |
| Related constraint | First-frame and reference images are mutually exclusive | MiniMax |
Inclusion rule. Statements bearing on more than one speaker in a single shot. Where nothing is published the row says so, because the absence is the finding on this page. Order. From direct guidance to adjacent statements.
1An implied decision is not a documented one
Tying lip-sync to whoever is on screen only makes sense if something decides who that is when there are two candidates. That decision exists in the implementation and is absent from the documentation, so a production cannot plan around it or prompt against it.
This is the clearest example in the register of documentation written for a single-subject case being read by people working on a multi-subject one.
2A constraint from elsewhere lands squarely on this shot
One model documents that a first-frame image and reference images cannot be used in the same call. That is not about audio at all, and it collides with exactly this case: continuing from a known frame and carrying two character references are the two things a conversation between established characters needs at once.
Recording it on a page about speech looks like a category error until the shot is considered. A register that kept constraints strictly inside their own topic would describe a workflow the documentation does not permit.
3Why this page exists with almost nothing in it
A register that only published populated rows would suggest this question has no answer because nobody asks it. The opposite is true: it is the commonest shot in the form and the least documented behaviour in the category.
The page will fill when a vendor writes two sentences about it. Until then it records that seventeen entries were read for it and one of them offers a sentence.
- Audio sourceNative speech with lip-sync tied to the speaker who is on screenspeaker-aware
- Constraint that collides with voice workA first-frame image and reference images cannot be used in the same call
- Two characters in frameFor multi-character scenes the guide asks for speakers labelled consistently and turns alternated, so each line lands on the right characteraddressed rather than skipped
- Dialogue timing against clip lengthA four-second shot is said to accommodate one or two short exchanges and an eight-second clip a few more, with long speeches unlikely to sync
- A framing condition attached to lip-syncLip sync performs best when the avatar is framed closer in the scene rather than positioned far from the camera
4Sources
Every line is read from the vendor documentation linked on the speech controls page, checked 2026-09-12. What counts as documented is on how read. Nearby: Who docs are for, How long a shot can be, Fifteen seconds of voice.