Writing dialogue a model can perform
A model reads exactly what is written, including the things a writer usually leaves to a performer. Sentence length sets the breath, punctuation sets the pauses, and anything abbreviated gets expanded by a rule nobody chose. As of 2026-09-12.
| In the text | Decides |
|---|---|
| Abbreviations and names | Whether the read is even intelligible |
| Numbers and dates | How they are spoken |
| Punctuation | Where the pauses fall |
| Sentence length | Whether the read hurries |
Inclusion rule. Properties of written dialogue that change generated speech. Order. Alphabetical by element.
1Length sets the breath
Long sentences produce a read that hurries, because the whole clause is planned as one unit. Short sentences produce space. Where a line has to sit inside a fixed shot length, the sentence structure is the main lever available.
Counting words against the shot length before generating is faster than regenerating a line that overruns by a second. Most languages settle around a predictable rate, and a production learns its own quickly.
2Punctuation is direction
A comma, a full stop, an ellipsis and a dash produce different pauses of different lengths. Writing dialogue with light punctuation and expecting the model to find the beats produces a flat, even read.
Conversely, heavy punctuation produces a mannered one. The useful setting is usually somewhere a writer would consider slightly over-punctuated on the page.
3Numbers, dates and symbols get expanded
A figure written as digits is spoken according to whatever rule the model applies, which may not be the one the scene needs. The same applies to dates, currency, times and units.
Writing them out in words removes the guess entirely. It reads oddly in the script and correctly in the audio, which is the trade worth making.
4Abbreviations and names are where reads break
Initialisms, brand names, foreign names and invented words are the reliable failure cases. There is no way to know how one will be read except by generating it and listening.
Test the proper nouns of a production before writing a season around them. Renaming a character in episode one is cheap; discovering in episode nine that the name is read wrong is not.
5Emphasis has to be in the words
Where a model offers no explicit emphasis control, the only lever is word order and sentence construction. Putting the important word at the end of a short sentence is more reliable than any markup that may or may not be honoured.
Where explicit controls do exist, they are documented per model and vary considerably, which is what the control table records.
6Write the pauses as lines
Silence between lines is part of the performance and is easier to control in the edit than in the text. Splitting an exchange into separate generations, with the gap decided afterwards, gives more control than trying to write the pause into one block.
It also means a single bad line can be regenerated without touching the rest of the scene, which is usually the larger saving.
7Interruptions and overlaps have to be faked
Two characters talking over each other is one of the things generated speech does not do, because each line is produced on its own. The overlap is built in the edit from two clean generations rather than asked for in the text.
Writing the lines so they can overlap — each complete on its own, with the collision decided later — is what makes that possible. Lines written as fragments of one interrupted sentence do not cut together.
8Read it aloud before generating it
Anything that is awkward to say is awkward to generate, and the model has no instinct for reworking it. Tongue-twisters, stacked clauses and repeated consonants all come back sounding like what they are.
Reading each line once out loud catches most of this at the writing stage, where a rewrite is free.
9Keep the script and the spoken text separate
The version a person reads and the version fed to the model diverge as soon as numbers are written out and punctuation is adjusted. Keeping both, with the spoken version stored beside the shot, is what makes a later regeneration reproduce the same read.
Losing the spoken version is the common reason a regenerated line does not match the one it replaces.
10What the table records instead
These are writing habits and they assert nothing about any model. Which controls each model documents — language lists, voice selection, emphasis, timing — is recorded in the control table with the page each line came from.
Use the table to find out what a given model lets you ask for; use this page to write the text you will be asking it to read.
11Where the published detail is logged
What is described above is how the techniques work in general. Which models document which of them, in whose words, is kept in the speech controls table with a date on every field.
- The control table — what each model documents, with sources
A mechanism note rather than a documented field. Nothing here is attributed to a model, and nothing here is a claim about one. The sourced material is on the speech controls table. Related: Dub or regenerate, How lip-sync works.