What happens to the voice when a shot is redone
A shot gets redone for visual reasons constantly. Where the sound was generated with it, the performance is regenerated too and an approved line is gone. Where the sound was supplied, the picture changes and the read stays. As of 2026-09-12.
| Model | What is documented | Detail |
|---|---|---|
| Generated with the picture | Eight entries | A re-render is a new performance |
| Supplied as a recording | Three entries | The read survives any re-render |
| Synthesised from a script | Four entries | Re-run the speech call, or leave it |
Inclusion rule. The three documented origins of the sound, with the number of entries on each. An entry publishing nothing about where its audio comes from is not placed in any of the three. Order. Fixed order: sound generated, sound supplied, sound synthesised separately.
1The cost nobody budgets for
An approved line disappearing because a background was wrong is an expense that appears in no rate card. It happens on every entry that makes sound with the picture, and it happens silently: the new file is valid, in sync, and different.
Productions that notice this separate picture fixes from line fixes wherever the tool allows it. Most of this register does not allow it, which is an argument for approving picture and audio in one pass rather than two.
2Supplied audio makes a re-render cheap and a re-record expensive
Handing over a recording moves the fragile part upstream. A visual fix costs one generation and leaves the performance untouched; a line change costs a room, a performer and a diary, which is a larger and much more predictable cost.
Which of the two a production prefers depends on what changes more often. Series with locked scripts and restless art direction want supplied audio; series still finding their dialogue want it generated.
3A separate speech call is the middle position
Where speech is its own endpoint, the audio is an artefact with a name. A picture fix leaves it alone, a line fix costs one cheap call, and the two never collide. That is the arrangement that behaves best under revision and it is documented by four entries here.
What it does not give is a performance. A synthesised read can be regenerated endlessly and will not become a different interpretation, because nothing in these products takes direction.
- Audio sourceNative audio on VIDEO 3.0 in five languagesgenerated with the picture
- Audio sourceAvatar videos are driven by audio, and the character in the image will lip-sync and move to the audio providedhanded to the model rather than made by it
- Audio sourceA text to speech endpoint turns a script into speech audio as a step of its ownan endpoint apart from the picture
- Audio sourceThe audio parameter defaults to true so the returned video carries an audio track, and false returns one withouton unless it is switched off
4Sources
Every line is read from the vendor documentation linked on the speech controls page, checked 2026-09-12. What counts as documented is on how read. Nearby: Consent in the docs, A second control, What kind of page.