Sentioscope

Speech and voice controls, as each vendor documents them

Timbre: the part of a voice a short clip carries

Timbre is what makes a voice recognisable independently of what it is saying. It is the property a fifteen-second clip carries, which is why the published caps in this register are as short as they are. As of 2026-09-12.

What a fifteen-second clip actually copiesTimbre is what makes a voice recognisable regardless of what it says, and it survives a very short sample. Accent, pacing and emotional register depend on the reading rather than on the voice.Survives a short clipDepends on the readingWhat it isTimbre: the colour of the voiceAccent, pacing, emotional registerWhat to chooseClean, ordinary speechNothing; do not choose for theseDocumented by anyoneNoNoWhere it showsImmediately, in one lineAcross a season, as driftNo vendor here says which properties a model attends to
Fig. 1 Because timbre carries and the rest does not, a rebuilt voice sounds like the same person behaving differently, which is invisible in one shot.
Timbre, as published. Recorded 2026-09-12.
PropertyCarried by a short clipWhat it depends on instead
AccentUnpredictablyHow much of the sample shares the accent
Emotional registerRarelyThe reading in the clip, not the voice in it
PacingRarelyThe prompt, or the length of the shot
TimbreReliablyAlmost nothing else

Inclusion rule. Properties of a spoken performance, sorted by how dependably each survives being copied from a short reference recording. No vendor in this register documents any of them. Order. Alphabetical by property.

1A clip is chosen for timbre, not for a good reading

Fifteen seconds is enough to establish what a voice sounds like and not enough to establish how it acts. A clean stretch of ordinary speech therefore outperforms a theatrical excerpt, because the excerpt had to be trimmed and the trim decides what the model hears.

The advice is counter-intuitive and produces steadier results across a season. Nothing here is documented by any vendor, which is why this page is a definition rather than a row in the table.

2Nobody publishes which properties a model attends to

Across the entries with a published sample window, not one says whether accent, pacing or emotional colour are copied. A production supplying a clip is therefore making a bet on an undocumented behaviour, and the bet is cheapest when the clip is unremarkable.

That gap is the reason a preview matters so much. Hearing the resulting voice settles in one listen what no page in this register explains.

3Timbre is also what makes drift invisible

Because timbre carries and the rest does not, a voice rebuilt from a slightly different clip usually sounds like the same person behaving differently. That is hard to notice in one shot and obvious across a season.

Which is why the clip belongs in the archive next to the visual reference. A regenerated shot months later with whatever file was at hand produces a performance shift nothing in the output announces.

4Sources read for this entry

This page defines the term and logs what each vendor documents about it, each figure read on 2026-09-12. Related: Digital replica, Clip length.