Sentioscope

Speech and voice controls, as each vendor documents them

Writing dialogue a model can perform

A model reads exactly what is written, including the things a writer usually leaves to a performer. Sentence length sets the breath, punctuation sets the pauses, and anything abbreviated gets expanded by a rule nobody chose. As of 2026-09-12.

What a model reads in a line that a performer would have decidedA model reads exactly what is on the page, including decisions a writer normally leaves to a performer. Sentence length sets where the breath falls. Punctuation sets the pauses. Numbers, dates, symbols and abbreviations are expanded by a rule nobody on the production chose, and emphasis exists only if it is in the words.What a performer decidesWhat the written line decidesBreathWhere to take oneSentence lengthPausesWhere to holdPunctuationNumbers and symbolsHow to say themA rule nobody on the production choseEmphasisWhich word to lean onOnly what the words themselves carryKeep the script and the spoken text as separate documents
Fig. 1 Everything a performer would have supplied has to be written down, and the parts nobody writes down are the parts that come back wrong.
What the text decides about the read. Recorded 2026-09-12.
In the textDecides
Abbreviations and namesWhether the read is even intelligible
Numbers and datesHow they are spoken
PunctuationWhere the pauses fall
Sentence lengthWhether the read hurries

Inclusion rule. Properties of written dialogue that change generated speech. Order. Alphabetical by element.

1Length sets the breath

Long sentences produce a read that hurries, because the whole clause is planned as one unit. Short sentences produce space. Where a line has to sit inside a fixed shot length, the sentence structure is the main lever available.

Counting words against the shot length before generating is faster than regenerating a line that overruns by a second. Most languages settle around a predictable rate, and a production learns its own quickly.

2Punctuation is direction

A comma, a full stop, an ellipsis and a dash produce different pauses of different lengths. Writing dialogue with light punctuation and expecting the model to find the beats produces a flat, even read.

Conversely, heavy punctuation produces a mannered one. The useful setting is usually somewhere a writer would consider slightly over-punctuated on the page.

3Numbers, dates and symbols get expanded

A figure written as digits is spoken according to whatever rule the model applies, which may not be the one the scene needs. The same applies to dates, currency, times and units.

Writing them out in words removes the guess entirely. It reads oddly in the script and correctly in the audio, which is the trade worth making.

4Abbreviations and names are where reads break

Initialisms, brand names, foreign names and invented words are the reliable failure cases. There is no way to know how one will be read except by generating it and listening.

Test the proper nouns of a production before writing a season around them. Renaming a character in episode one is cheap; discovering in episode nine that the name is read wrong is not.

5Emphasis has to be in the words

Where a model offers no explicit emphasis control, the only lever is word order and sentence construction. Putting the important word at the end of a short sentence is more reliable than any markup that may or may not be honoured.

Where explicit controls do exist, they are documented per model and vary considerably, which is what the control table records.

6Write the pauses as lines

Silence between lines is part of the performance and is easier to control in the edit than in the text. Splitting an exchange into separate generations, with the gap decided afterwards, gives more control than trying to write the pause into one block.

It also means a single bad line can be regenerated without touching the rest of the scene, which is usually the larger saving.

7Interruptions and overlaps have to be faked

Two characters talking over each other is one of the things generated speech does not do, because each line is produced on its own. The overlap is built in the edit from two clean generations rather than asked for in the text.

Writing the lines so they can overlap — each complete on its own, with the collision decided later — is what makes that possible. Lines written as fragments of one interrupted sentence do not cut together.

8Read it aloud before generating it

Anything that is awkward to say is awkward to generate, and the model has no instinct for reworking it. Tongue-twisters, stacked clauses and repeated consonants all come back sounding like what they are.

Reading each line once out loud catches most of this at the writing stage, where a rewrite is free.

9Keep the script and the spoken text separate

The version a person reads and the version fed to the model diverge as soon as numbers are written out and punctuation is adjusted. Keeping both, with the spoken version stored beside the shot, is what makes a later regeneration reproduce the same read.

Losing the spoken version is the common reason a regenerated line does not match the one it replaces.

10What the table records instead

These are writing habits and they assert nothing about any model. Which controls each model documents — language lists, voice selection, emphasis, timing — is recorded in the control table with the page each line came from.

Use the table to find out what a given model lets you ask for; use this page to write the text you will be asking it to read.

11Where the published detail is logged

What is described above is how the techniques work in general. Which models document which of them, in whose words, is kept in the speech controls table with a date on every field.

A mechanism note rather than a documented field. Nothing here is attributed to a model, and nothing here is a claim about one. The sourced material is on the speech controls table. Related: Dub or regenerate, How lip-sync works.