Sentioscope

Speech and voice controls, as each vendor documents them

Casting a voice you can keep for forty episodes

Picking a voice is the easy half. The half that matters is whether the same voice comes back in episode forty, generated by a different person from a different machine, and that depends on where the voice is stored rather than on how it sounds. As of 2026-09-12.

Where a cast voice is kept, and whether it comes back in episode fortyThree places a voice can live. A preset chosen from a library is identified by a name anyone on the team can select again. A supplied recording lives in your own files and comes back only if somebody kept the clip. Speech generated with the picture is not stored anywhere between shots, so continuity depends on the prompt being identical.Most retrievable firstA preset libraryNamed, finite, and selectable by anybody on the team. The samename returns the same voice.A supplied recordingLives in your files. It survives only as long as the clip iskept with the character.Generated with the pictureNothing is stored between shots. Continuity rests on the requestbeing made the same way again.Episode forty is generated by somebody else on a different machine
Fig. 1 Casting sounds like a listening decision and is really a storage decision. The question is not which voice you like, it is which one a different person can retrieve later.
What a voice choice has to survive. Recorded 2026-09-12.
ConstraintWhy it binds
Across episodesThe same character returns for the whole run
Across languagesA market version needs a counterpart voice
Across regenerationsA redone shot must match the approved one

Inclusion rule. Constraints a voice decision has to satisfy in serial work. One-off narration is a different case. Order. Alphabetical by constraint.

1Three places a voice can come from

A preset library, where you choose from voices the vendor provides. A supplied recording, where you give a sample and the system matches it. Or whatever the model produces unguided, which is not casting so much as accepting.

The first is reproducible by identifier and the least expressive. The second gives more control and introduces a dependency on a file somebody has to keep. The third is fine for one clip and unusable for a series.

2The fifteen-second problem

Where a system takes reference audio, the cap is often very short: a handful of seconds across a couple of clips. That is a tiny amount of material from which to characterise a voice, and it puts unusual weight on which seconds you choose.

Practical advice that follows: pick a sample with normal speaking pace, no music, no room echo, and a range of pitch rather than a monotone line. A dramatic delivery makes a striking sample and an unusable reference, because everything generated from it inherits the drama.

3Storage is the whole ball game for a series

A voice attached to a reusable character object comes back automatically. A voice re-established from a file on every call comes back only if somebody attaches the right file. Over forty episodes and several people, the difference is not subtle.

If the system does not store voices, then your reference clips are production assets with the same status as the character sheet: versioned, named, kept with the project, and backed up.

4Cloning a real person is a different decision

Supplying a recording of an identifiable person, without that person's agreement, is a likeness question rather than a technical one. Some vendors say so explicitly on their own pages and rule it out; most say nothing.

This site records who has published a boundary and does not offer a view on the law. What is worth noting operationally is that the decision gets made at the moment somebody uploads a file, usually by whoever is quickest, which is an argument for settling it before production rather than during.

5Language is part of casting, not a setting applied later

If the series ships in several markets, the voice question multiplies: the same character needs a voice in each language, and matching a preset across languages is harder than matching one within a language. A named list of supported languages is therefore a casting constraint, not a marketing number.

Where a vendor publishes a count rather than names, a producer targeting a specific market has no answer at all. That is the practical reason this register refuses to treat a count and a list as comparable.

6If the series ships in several languages, start from the list

The constraint that binds hardest on a multi-market series is which languages the tool will actually speak, and a named list is the only form of that answer a producer can act on. Among the models tracked here, SceneMixer publishes 15 named dialogue languages with Cantonese available for dialogue only, and states that the voice comes from a preset library or the user's own recordings.

Kling AI publishes native audio in five languages without naming them, which answers nothing about any particular market however the number is read. The other models in this register publish neither a list nor a count. Whether a named list matters more to you than the other fields in the table is a judgement about your show; that only one entry here supplies one is a fact about the published material.

7Listen to the whole cast together before locking

Voices chosen one at a time tend to cluster: two characters end up in the same register and an audience stops distinguishing them by ear. Generating one line per main character and listening to them back to back catches this in ten minutes and is almost never done.

Do it again per language. A cast that separates cleanly in one language can collapse into sameness in another, because the preset libraries are not evenly deep across languages.

8A short casting protocol

Choose the voice before the first episode, not during. Record the identifier or keep the reference file with the character sheet. Generate one line per main character in every language you will ship, and listen to them together. Then lock it, and treat a change as a recast rather than a setting tweak.

That last framing does most of the work. Teams that treat voice as a parameter change it casually; teams that treat it as casting do not.

9Where the published detail is logged

What is described above is how the techniques work in general. Which models document which of them, in whose words, is kept in the speech controls table with a date on every field.

  • SceneMixer — names a preset library and user recordings as the two routes, and publishes 15 named dialogue languages
  • MiniMax — caps supplied reference audio at 15 seconds across at most three clips

A mechanism note rather than a documented field. Nothing here is attributed to a model, and nothing here is a claim about one. The sourced material is on the speech controls table. Related: Three routes to a voice, Reading a blank cell.