Sentioscope

Speech and voice controls, as each vendor documents them

On-screen speaker: which face the sound belongs to

The on-screen speaker is the character a line is attributed to inside a shot. Four entries address the question and no two address it alike: a model behaviour, an identifier on a request, labels in a prompt, and a framing condition. As of 2026-09-12.

Four answers to whose mouth should moveA model behaviour with no override, an identifier on a request, labels inside a prompt, and a framing condition. Thirteen entries in this register say nothing about the question at all.Two people in frame, one of them talkingA documented behaviourThe model picksNo parameter reaches it,so coverage is the onlyleverAn id or a labelYou pickControllable, and onlyinside one generationNothing publishedAlternating singlesTwo generations perexchange, and a consistentresultOne entry names framing as the condition, which is a camera decision
Fig. 1 Where nothing is published, the safe assumption is alternating singles, which doubles what a conversation costs to generate.
On-screen speaker, as published. Recorded 2026-09-12.
ModelHow the speaker is decidedWhat a production controls
MiniMaxThe model follows whoever is on screenOnly the framing
PixVerseA speaker id passed on the requestWhich voice, and one part per call
Sora 2Labels in the prompt, turns alternatedWhich line lands on which face
SynthesiaFraming is said to affect performanceHow close the avatar sits to camera

Inclusion rule. Entries whose documentation says anything about attributing speech to a particular face. An entry that mentions lip-sync without addressing who is speaking does not earn a row. Order. Alphabetical by model name.

1The commonest shot in drama is barely addressed

Two people, one of them talking, is the staple of dialogue coverage, and thirteen entries in this register say nothing about it. The four that do between them offer a behaviour, a routing field, a prompt convention and a framing hint.

Where nothing is published, the safe assumption is alternating singles. That doubles the generations a conversation costs, and it is a planning decision rather than a preference.

2A behaviour cannot be overridden and an instruction can

Where the model is documented as following whoever is on screen, no parameter reaches the decision, so coverage is the only lever: a single with the other voice off-frame removes the ambiguity entirely.

Where labels route lines inside a prompt, the writer decides. That is more controllable, and it still only holds inside one generation, because nothing connects a label to a voice that survives the next call.

3A framing condition names the weak case

Saying lip sync performs better with the avatar framed closer is a quiet statement that the wide two-shot is where alignment struggles. Naming a weak case constrains a shot list in a way that a capability claim never does.

It is also the only sentence here that connects this question to a camera decision rather than to a parameter, which is where the answer usually has to be found anyway.

4Sources read for this entry

This page defines the term and logs what each vendor documents about it, each figure read on 2026-09-12. Related: Native audio, Reference audio.