A line that fits a shot can still be unreadable
Sizing a line to a shot is one constraint. Sizing it to a reading speed is another, and the tighter one wins. For a format watched muted as often as not, the subtitle is the channel that has to work. As of 2026-09-12.
| Ceiling | Measured in | Who it binds |
|---|---|---|
| A shot length | Seconds of generated footage | Whoever writes the line |
| A subtitle reading speed | Characters a second | Whoever reads it |
| A translated line | Characters, after expansion | Whichever market expands most |
Inclusion rule. Constraints that apply to one line of dialogue simultaneously, each measured differently. Order. Alphabetical by ceiling.
1The two ceilings are not measured in the same units
A shot gives seconds and a subtitle gives characters a second, so a line can sit comfortably inside one and exceed the other. Published reading limits are also lower for younger audiences, which matters for a format distributed by app.
Doing both sums while writing is the only cheap moment. Afterwards one of the two has to be violated, and whichever is chosen the audience notices.
2Translation expands text, unevenly
A line sized for a shot in one language will not be sized for it in another, because the same meaning takes a different number of characters. Markets that expand most are the ones where subtitles overrun first.
Where a production generates in each language rather than translating afterwards, the spoken length changes too, and the shot it was written for may no longer hold it. Neither problem is documented by any vendor here.
3Muted viewing changes which channel is primary
Short-form distribution means a large share of viewing happens with the sound off. That makes the subtitle the performance for those viewers, and it makes reading speed a creative constraint rather than an accessibility one.
Which in turn changes how dialogue should be written: shorter lines, fewer subordinate clauses, and emphasis carried by word order rather than by delivery. None of that depends on which model is used.
4Where the published detail is logged
What is described above is how the techniques work in general. Which models document which of them, in whose words, is kept in the speech controls table with a date on every field.
- Dialogue timing — how many words fit a generated shot
- Clip length — the unit a line has to fit inside
A mechanism note rather than a documented field. Nothing here is attributed to a model, and nothing here is a claim about one. The sourced material is on the speech controls table. Related: Asking a vendor, Asking for silence.