Sentioscope

Speech and voice controls, as each vendor documents them

What happens with two speakers in frame

One vendor now addresses it: the Sora 2 prompting guide asks for speakers to be labelled consistently and turns alternated, so each line lands on the right character. Two others say something adjacent. Everybody else is silent on the shot a dialogue scene is built from. As of 2026-09-12.

What the register found on the shot a dialogue scene is built fromOne vendor publishes guidance: label the speakers and alternate their turns so each line lands on the right character. One ties lip movement to whoever is on screen without saying how the choice is made. One warns that lip sync holds better framed closer, which makes the wide two-shot the weak case. Nobody else addresses it.Closest thing to an answerGuidance, from one guideSpeakers labelled consistently, turns alternated, so lines reachthe right face.An implied selectionLip movement follows the on-screen speaker; how that speaker ispicked is unstated.A framing warningLip sync performs better framed closer, which is where twopeople rarely fit.SilenceEvery other entry publishes nothing about two charactersspeaking in one frame.Where the register stops
Fig. 1 Three statements, none of them a mechanism, on a shot that appears in almost every scene of a drama.
What the register found on two speakers in one frame. Recorded 2026-09-12.
ModelWhat is documentedDetail
Documented guidanceLabel speakers, alternate turnsSora 2 prompting guide
Nearest mechanismLip-sync tied to the on-screen speakerMiniMax video guide
A framing conditionLip sync performs better framed closerSynthesia avatar docs
Related constraintFirst-frame and reference images are mutually exclusiveMiniMax

Inclusion rule. Statements bearing on more than one speaker in a single shot. Where nothing is published the row says so, because the absence is the finding on this page. Order. From direct guidance to adjacent statements.

1An implied decision is not a documented one

Tying lip-sync to whoever is on screen only makes sense if something decides who that is when there are two candidates. That decision exists in the implementation and is absent from the documentation, so a production cannot plan around it or prompt against it.

This is the clearest example in the register of documentation written for a single-subject case being read by people working on a multi-subject one.

2A constraint from elsewhere lands squarely on this shot

One model documents that a first-frame image and reference images cannot be used in the same call. That is not about audio at all, and it collides with exactly this case: continuing from a known frame and carrying two character references are the two things a conversation between established characters needs at once.

Recording it on a page about speech looks like a category error until the shot is considered. A register that kept constraints strictly inside their own topic would describe a workflow the documentation does not permit.

3Why this page exists with almost nothing in it

A register that only published populated rows would suggest this question has no answer because nobody asks it. The opposite is true: it is the commonest shot in the form and the least documented behaviour in the category.

The page will fill when a vendor writes two sentences about it. Until then it records that seventeen entries were read for it and one of them offers a sentence.

  • Audio source
    Native speech with lip-sync tied to the speaker who is on screenspeaker-awareMiniMax, video generation guide / recorded 2026-09-12
  • Constraint that collides with voice work
    A first-frame image and reference images cannot be used in the same callMiniMax, video generation guide / recorded 2026-09-12
  • Two characters in frame
    For multi-character scenes the guide asks for speakers labelled consistently and turns alternated, so each line lands on the right characteraddressed rather than skippedOpenAI, Sora 2 prompting guide / recorded 2026-09-22
  • Dialogue timing against clip length
    A four-second shot is said to accommodate one or two short exchanges and an eight-second clip a few more, with long speeches unlikely to syncOpenAI, Sora 2 prompting guide / recorded 2026-09-22
  • A framing condition attached to lip-sync
    Lip sync performs best when the avatar is framed closer in the scene rather than positioned far from the cameraSynthesia, create an avatar / recorded 2026-09-22

4Sources

Every line is read from the vendor documentation linked on the speech controls page, checked 2026-09-12. What counts as documented is on how read. Nearby: Who docs are for, How long a shot can be, Fifteen seconds of voice.