Sentioscope

Speech and voice controls, as each vendor documents them

Separating a picture revision from an audio one

Shots get redone for visual reasons constantly. Whether that costs a performance depends entirely on whether the audio came from the same call, and most entries in this register give no way to separate the two. As of 2026-09-12.

Three architectures, and what a revision costs on eachWhere sound and picture come from one call, a visual fix returns a new performance. Where speech is its own call, the two never collide. Where audio was supplied, a line change is a session.A picture fix costsA line fix costsA separate speech callThe render onlyOne cheap speech callSound with the pictureA new performance as wellA generationSound suppliedThe render onlyA room and a performerWhich to preferIf art direction moves mostIf dialogue moves mostAn approved line vanishing appears on no rate card
Fig. 1 The choice is effectively irreversible once a season is in production, because changing architecture means re-establishing every voice.
What each architecture does to an approved line when a shot is generated again. Recorded 2026-09-12.
ArchitectureA picture fix costsA line fix costs
A separate speech callThe render onlyOne cheap speech call
Sound generated with the pictureA new performance as wellA generation
Sound supplied as a recordingThe render onlyA room and a performer

Inclusion rule. The three documented arrangements, compared by what a revision costs on each. Order. Alphabetical by architecture.

1The expense nobody puts in a plan

An approved line vanishing because a background was wrong appears on no rate card. The replacement file is valid, aligned and different, so there is nothing to diagnose and nothing to appeal to.

Productions that have been through it approve picture and audio in one pass rather than two, which is slower and avoids re-opening a decision that was already made.

2A separate speech stage survives revision most easily

Where speech comes from its own call, the audio is an artefact with a name. A picture fix leaves it alone, a word fix costs one cheap call, and the two never collide. That is the arrangement to prefer if a script is still moving.

What it does not give is a performance. A synthesised read can be regenerated endlessly and will not become a different interpretation, because nothing in these products takes direction.

3Supplied audio trades one cost for a bigger, more predictable one

Handing over a recording makes visual revision free and line revision expensive. For a series with a locked script and restless art direction that is the right way round; for one still finding its dialogue it is exactly wrong.

The choice is worth making deliberately at the start, because it is effectively irreversible once a season is in production. Changing architecture mid-run means re-establishing every voice.

4Where the published detail is logged

What is described above is how the techniques work in general. Which models document which of them, in whose words, is kept in the speech controls table with a date on every field.

A mechanism note rather than a documented field. Nothing here is attributed to a model, and nothing here is a claim about one. The sourced material is on the speech controls table. Related: Framing and the mouth, Auditioning first.