How framing decides whether alignment holds up
Alignment is usually discussed as a model property. One vendor attaches a condition to it instead: performance improves with the subject closer to camera. That makes the weakest case the widest shot, which is also the staple of dialogue coverage. As of 2026-09-12.
| Shot size | What the mouth has to sell | How forgiving it is |
|---|---|---|
| A close-up | Every consonant, in detail | Least forgiving of detail, most of timing |
| A medium | Timing, and rough shapes | The usual working compromise |
| A wide two-shot | Which of two people is talking | Least forgiving of all |
Inclusion rule. Standard shot sizes, considered against the alignment problem rather than against composition. Order. Alphabetical by shot size.
1A close-up is harder in detail and easier in ambiguity
Filling the frame with a face means every shape has to be right, and it also means there is only one candidate. The difficulty moves from deciding whose mouth to move onto making one mouth convincing, which is the problem models are actually improving at.
That is why close coverage is the safe default for generated dialogue even though it looks like the demanding choice. The demanding part is the part being handled.
2A wide two-shot asks for a decision nobody documented
Two faces at a distance combine the least detail with the most ambiguity. A model has to pick a speaker, and at that size an error is legible without being obviously wrong, which is the worst combination for a reviewer.
Almost nothing in this register addresses the case, so a production planning wide coverage for conversation is planning on an undocumented behaviour. That is a reasonable thing to test and a poor thing to assume.
3Coverage is a lever that exists everywhere
Unlike a parameter, framing is available on every product here. Where a vendor documents a behaviour rather than a control, changing the shot is the only way to change the outcome, and it is a change a director was going to make anyway.
Writing that into a shot list at the start is cheaper than discovering it. A production that covers conversation in singles by convention never has to find out what a given model does with a two-shot.
4Where the published detail is logged
What is described above is how the techniques work in general. Which models document which of them, in whose words, is kept in the speech controls table with a date on every field.
- Covering a conversation — and what each option costs
- The lip-sync column — what each vendor says moves the mouth
A mechanism note rather than a documented field. Nothing here is attributed to a model, and nothing here is a claim about one. The sourced material is on the speech controls table. Related: Auditioning first, Singing and read copy.