Sentioscope

Speech and voice controls, as each vendor documents them

Covering a conversation when nothing promises the mouth

A two-hander asks a model to decide which of two faces is speaking. Where nothing is published about that decision, the reliable answer is to remove it: shoot singles and keep the other voice out of frame. As of 2026-09-12.

Removing the choice instead of prompting around itIf only one face is in frame, no decision exists for a model to get wrong. The instinct when a two-shot comes back wrong is to change the prompt, and the prompt is usually not where the fix lives.Write the exchangeTurns labelled, andeach line shortenough to fit a shot.Then coverCover in singlesEach line on thespeaker, with theother voiceoff-frame.No choice leftGenerateTwo calls perexchange, and aresult that does notdepend on a guess.
Fig. 1 An exchange covered in singles costs one call per line rather than one per exchange, which becomes a budget line across a series.
Ways to cover a conversation, and what each one costs. Recorded 2026-09-12.
CoverageRemoves the ambiguityWhat it costs
A single, other voice off-frameCompletelyTwo generations per exchange
An over-shoulderMostly, if the listener faces awayOne generation, harder to compose
A two-shot, one speakingNot at allOne generation, and a risk
A wide with both facesLeast of allOne generation, and the weakest case

Inclusion rule. Standard coverage options for a dialogue exchange, judged by whether they leave a model a choice to make. Order. Fixed order, from the coverage that removes the choice to the one that most invites it.

1Ambiguity is a composition problem before it is a model problem

If only one face is in frame, no decision exists. That is worth saying plainly, because the instinct when a two-shot comes back wrong is to change the prompt, and the prompt is usually not where the fix lives.

The cost is generations. An exchange covered in singles costs one call per line rather than one per exchange, and for a series that adds up into a real budget line rather than an inconvenience.

2One published condition points the same way

A vendor stating that alignment performs better with a figure framed closer to camera has said, quietly, that the wide two-shot is the weak case. That is a coverage instruction disguised as a tip, and it happens to agree with what singles imply.

Where a model is documented as following whoever is on screen, the same logic holds from the other direction: the decision it is making is one a production can avoid asking for.

3Labels help inside a call and not across one

Where a prompting convention exists for naming speakers and alternating turns, an exchange can be routed correctly inside one generation. That is real and it is bounded: nothing connects a label to a voice that survives the next call.

So labels solve the shot and coverage solves the scene. A production using both gets exchanges that play and characters that sound the same, and neither technique delivers the other.

4Where the published detail is logged

What is described above is how the techniques work in general. Which models document which of them, in whose words, is kept in the speech controls table with a date on every field.

A mechanism note rather than a documented field. Nothing here is attributed to a model, and nothing here is a claim about one. The sourced material is on the speech controls table. Related: Keeping the clip, Scheduling a track.