Filing the clip that made a character sound right
A reference clip is an asset, not a step. Whatever audio produced a character has to live beside that character under a name that says so, because a shot regenerated later with a different file returns a different person. As of 2026-09-12.
| What to record | Why | Where it usually goes wrong |
|---|---|---|
| The clip itself | It is the only copy of the voice | Left in a download folder |
| The exact trim | The trim decides what the model heard | Re-trimmed later, slightly differently |
| The character it belongs to | Nothing in the output records it | Named after the file, not the part |
| The date it was chosen | To explain a later shift | Never recorded at all |
Inclusion rule. Items a production has to keep itself where a vendor documents no stored voice. Order. Fixed order: the clip, its trim, the part it belongs to, then the date.
1The drift is silent, which is what makes it expensive
A voice rebuilt from a slightly different clip usually sounds like the same person behaving differently. No single shot looks wrong, and an audience comparing episode one with episode nine hears it immediately.
Because nothing in a returned file records which sample was used, there is no way to diagnose it after the fact. The only defence is that the file was never in doubt, which is an archival property rather than a technical one.
2Name the clip after the part, not after the take
A file called interview-final-2 will be reused by somebody who does not know which character it belongs to. A file named after the part is self-documenting, and it survives the person who chose it leaving the production.
Storing it beside the visual reference for the same character is the other half. Cast continuity is one problem, and splitting the picture half from the sound half across two systems guarantees they drift apart.
3Choose an unremarkable clip, and do it once
A clean stretch of ordinary speech is a more stable target than a dramatic excerpt that had to be cut, because the trim is what a model actually hears. That advice is counter-intuitive and it produces steadier results across a season.
Choosing once also matters. Re-auditioning a character mid-season because a better clip turned up is how a cast changes voice, and the improvement is rarely worth the discontinuity.
4Where the published detail is logged
What is described above is how the techniques work in general. Which models document which of them, in whose words, is kept in the speech controls table with a date on every field.
- Timbre — the part of a voice a short clip actually carries
- Fifteen seconds of voice — the published windows, compared
A mechanism note rather than a documented field. Nothing here is attributed to a model, and nothing here is a claim about one. The sourced material is on the speech controls table. Related: Scheduling a track, Open weights.