Covering a conversation when nothing promises the mouth
A two-hander asks a model to decide which of two faces is speaking. Where nothing is published about that decision, the reliable answer is to remove it: shoot singles and keep the other voice out of frame. As of 2026-09-12.
| Coverage | Removes the ambiguity | What it costs |
|---|---|---|
| A single, other voice off-frame | Completely | Two generations per exchange |
| An over-shoulder | Mostly, if the listener faces away | One generation, harder to compose |
| A two-shot, one speaking | Not at all | One generation, and a risk |
| A wide with both faces | Least of all | One generation, and the weakest case |
Inclusion rule. Standard coverage options for a dialogue exchange, judged by whether they leave a model a choice to make. Order. Fixed order, from the coverage that removes the choice to the one that most invites it.
1Ambiguity is a composition problem before it is a model problem
If only one face is in frame, no decision exists. That is worth saying plainly, because the instinct when a two-shot comes back wrong is to change the prompt, and the prompt is usually not where the fix lives.
The cost is generations. An exchange covered in singles costs one call per line rather than one per exchange, and for a series that adds up into a real budget line rather than an inconvenience.
2One published condition points the same way
A vendor stating that alignment performs better with a figure framed closer to camera has said, quietly, that the wide two-shot is the weak case. That is a coverage instruction disguised as a tip, and it happens to agree with what singles imply.
Where a model is documented as following whoever is on screen, the same logic holds from the other direction: the decision it is making is one a production can avoid asking for.
3Labels help inside a call and not across one
Where a prompting convention exists for naming speakers and alternating turns, an exchange can be routed correctly inside one generation. That is real and it is bounded: nothing connects a label to a voice that survives the next call.
So labels solve the shot and coverage solves the scene. A production using both gets exchanges that play and characters that sound the same, and neither technique delivers the other.
4Where the published detail is logged
What is described above is how the techniques work in general. Which models document which of them, in whose words, is kept in the speech controls table with a date on every field.
- Two speakers in frame — what the documentation says about it
- On-screen speaker — four answers to whose mouth moves
A mechanism note rather than a documented field. Nothing here is attributed to a model, and nothing here is a claim about one. The sourced material is on the speech controls table. Related: Keeping the clip, Scheduling a track.