Sentioscope

Speech and voice controls, as each vendor documents them

Supplying a voice instead of describing one

Reference audio is a clip supplied so the generated speech resembles it. Where it is documented the caps are published in seconds, and they range from fifteen seconds in total to a twenty-minute session, which makes the choice of clip a casting decision rather than a technical one. As of 2026-09-12.

What a short reference clip can and cannot carryReference audio is a clip supplied so that generated speech resembles it. Where the caps are documented they are low and the total duration allowed is short, which makes choosing the clip a casting decision. A clip carries timbre and delivery; it does not carry a language, a script or continuity between calls.What the clip carriesWhat it does notSoundTimbre and delivery of the voiceThe language the line will be inScopeThis callAnything stored for the next oneCastingWhich voice this character hasThat anybody else will pick the same clipKeep the clip filed with the character, not with the shot
Fig. 1 The cap is per call rather than per project, so the same short clip has to be good enough to be supplied again every time.
Reference audio, as published. Recorded 2026-09-12.
ModelClips per callLength and format
HeyGenOne recording, or a booked sessionInstant from one recording; professional from 20 minutes or more
MiniMax H3Up to three2 to 15 s per clip, 15 s in total, WAV or MP3
RunwayOne sample per voice10 s to 5 min, at most 10 MB, or a description of 20 characters
Wan 2.7 reference-to-videoOne voice per reference image or video1 to 10 s, WAV or MP3 up to 15 MB
Wan 3.0Up to five15 s in total, 1 to 15 s a clip, up to 15 MB

Inclusion rule. Models whose vendor publishes limits for audio supplied as the source of a voice. A model offering voice presets without accepting a recording is not given a row. Order. Alphabetical by model name.

1Fifteen seconds is the real constraint

Two vendors cap the total reference audio at fifteen seconds regardless of how many clips it is split across. That is enough to carry a voice and not enough to carry a performance, so the clip should be chosen for timbre rather than for a good reading.

A clean fifteen seconds of ordinary speech outperforms a dramatic thirty-second excerpt that had to be trimmed, because the trim decides what the model hears.

2One voice per reference asset changes the casting

Where a voice attaches to a specific reference image, each character carries its own audio and the cap applies per asset rather than per call. That is a different model of casting from a pool of voices, and it scales differently across a season.

3Keep the clip with the character

Whatever clip produced a voice has to be stored beside the character it belongs to, in the same place as the visual reference. A regenerated shot months later with a different clip produces a different person, and nothing in the output announces it.

4What a reference clip cannot carry

Timbre transfers. Accent, pacing and emotional register transfer unpredictably, and none of the vendors here document which of those the model attends to. A clip chosen for a dramatic reading can therefore produce a voice that sounds right and performs differently.

Choosing an unremarkable clip is counter-intuitive and produces steadier results across a season.

5The cap is per call, not per project

Fifteen seconds in total applies to each generation, so a scene with three speaking characters splits that budget three ways. For a conversation, the practical consequence is that voices are usually supplied one per shot rather than all at once.

That in turn means the shot list has to record which voice belongs to which shot, because the call cannot hold the whole cast.

6Sources read for this entry

This page defines the term and logs what each vendor documents about it, each figure read on 2026-09-12. Related: Dialogue timing, Lip-sync.