On-screen speaker: which face the sound belongs to
The on-screen speaker is the character a line is attributed to inside a shot. Four entries address the question and no two address it alike: a model behaviour, an identifier on a request, labels in a prompt, and a framing condition. As of 2026-09-12.
| Model | How the speaker is decided | What a production controls |
|---|---|---|
| MiniMax | The model follows whoever is on screen | Only the framing |
| PixVerse | A speaker id passed on the request | Which voice, and one part per call |
| Sora 2 | Labels in the prompt, turns alternated | Which line lands on which face |
| Synthesia | Framing is said to affect performance | How close the avatar sits to camera |
Inclusion rule. Entries whose documentation says anything about attributing speech to a particular face. An entry that mentions lip-sync without addressing who is speaking does not earn a row. Order. Alphabetical by model name.
1The commonest shot in drama is barely addressed
Two people, one of them talking, is the staple of dialogue coverage, and thirteen entries in this register say nothing about it. The four that do between them offer a behaviour, a routing field, a prompt convention and a framing hint.
Where nothing is published, the safe assumption is alternating singles. That doubles the generations a conversation costs, and it is a planning decision rather than a preference.
2A behaviour cannot be overridden and an instruction can
Where the model is documented as following whoever is on screen, no parameter reaches the decision, so coverage is the only lever: a single with the other voice off-frame removes the ambiguity entirely.
Where labels route lines inside a prompt, the writer decides. That is more controllable, and it still only holds inside one generation, because nothing connects a label to a voice that survives the next call.
3A framing condition names the weak case
Saying lip sync performs better with the avatar framed closer is a quiet statement that the wide two-shot is where alignment struggles. Naming a weak case constrains a shot list in a way that a capability claim never does.
It is also the only sentence here that connects this question to a camera decision rather than to a parameter, which is where the answer usually has to be found anyway.
4Sources read for this entry
This page defines the term and logs what each vendor documents about it, each figure read on 2026-09-12. Related: Native audio, Reference audio.