Sentioscope

Speech and voice controls, as each vendor documents them

How framing decides whether alignment holds up

Alignment is usually discussed as a model property. One vendor attaches a condition to it instead: performance improves with the subject closer to camera. That makes the weakest case the widest shot, which is also the staple of dialogue coverage. As of 2026-09-12.

A close-up is harder in detail and easier in ambiguityFilling the frame with one face means every shape has to be right and there is only one candidate. A wide two-shot combines the least detail with the most ambiguity, which is the worst pairing.Easiest for alignment to surviveA close-upEvery consonant visible, and only one mouth it could be.A mediumTiming and rough shapes; the usual working compromise.An over-shoulderOne face away from camera, which removes most of the choice.A wide two-shotLeast detail, most ambiguity, and almost nothing published aboutit.Hardest
Fig. 1 Close coverage is the safe default for generated dialogue even though it looks like the demanding choice, because the demanding part is the part being handled.
How the difficulty of a mouth changes with the size of the shot. Recorded 2026-09-12.
Shot sizeWhat the mouth has to sellHow forgiving it is
A close-upEvery consonant, in detailLeast forgiving of detail, most of timing
A mediumTiming, and rough shapesThe usual working compromise
A wide two-shotWhich of two people is talkingLeast forgiving of all

Inclusion rule. Standard shot sizes, considered against the alignment problem rather than against composition. Order. Alphabetical by shot size.

1A close-up is harder in detail and easier in ambiguity

Filling the frame with a face means every shape has to be right, and it also means there is only one candidate. The difficulty moves from deciding whose mouth to move onto making one mouth convincing, which is the problem models are actually improving at.

That is why close coverage is the safe default for generated dialogue even though it looks like the demanding choice. The demanding part is the part being handled.

2A wide two-shot asks for a decision nobody documented

Two faces at a distance combine the least detail with the most ambiguity. A model has to pick a speaker, and at that size an error is legible without being obviously wrong, which is the worst combination for a reviewer.

Almost nothing in this register addresses the case, so a production planning wide coverage for conversation is planning on an undocumented behaviour. That is a reasonable thing to test and a poor thing to assume.

3Coverage is a lever that exists everywhere

Unlike a parameter, framing is available on every product here. Where a vendor documents a behaviour rather than a control, changing the shot is the only way to change the outcome, and it is a change a director was going to make anyway.

Writing that into a shot list at the start is cheaper than discovering it. A production that covers conversation in singles by convention never has to find out what a given model does with a two-shot.

4Where the published detail is logged

What is described above is how the techniques work in general. Which models document which of them, in whose words, is kept in the speech controls table with a date on every field.

A mechanism note rather than a documented field. Nothing here is attributed to a model, and nothing here is a claim about one. The sourced material is on the speech controls table. Related: Auditioning first, Singing and read copy.