Supplying a voice instead of describing one
Reference audio is a clip supplied so the generated speech resembles it. Where it is documented the caps are published in seconds, and they range from fifteen seconds in total to a twenty-minute session, which makes the choice of clip a casting decision rather than a technical one. As of 2026-09-12.
| Model | Clips per call | Length and format |
|---|---|---|
| HeyGen | One recording, or a booked session | Instant from one recording; professional from 20 minutes or more |
| MiniMax H3 | Up to three | 2 to 15 s per clip, 15 s in total, WAV or MP3 |
| Runway | One sample per voice | 10 s to 5 min, at most 10 MB, or a description of 20 characters |
| Wan 2.7 reference-to-video | One voice per reference image or video | 1 to 10 s, WAV or MP3 up to 15 MB |
| Wan 3.0 | Up to five | 15 s in total, 1 to 15 s a clip, up to 15 MB |
Inclusion rule. Models whose vendor publishes limits for audio supplied as the source of a voice. A model offering voice presets without accepting a recording is not given a row. Order. Alphabetical by model name.
1Fifteen seconds is the real constraint
Two vendors cap the total reference audio at fifteen seconds regardless of how many clips it is split across. That is enough to carry a voice and not enough to carry a performance, so the clip should be chosen for timbre rather than for a good reading.
A clean fifteen seconds of ordinary speech outperforms a dramatic thirty-second excerpt that had to be trimmed, because the trim decides what the model hears.
2One voice per reference asset changes the casting
Where a voice attaches to a specific reference image, each character carries its own audio and the cap applies per asset rather than per call. That is a different model of casting from a pool of voices, and it scales differently across a season.
3Keep the clip with the character
Whatever clip produced a voice has to be stored beside the character it belongs to, in the same place as the visual reference. A regenerated shot months later with a different clip produces a different person, and nothing in the output announces it.
4What a reference clip cannot carry
Timbre transfers. Accent, pacing and emotional register transfer unpredictably, and none of the vendors here document which of those the model attends to. A clip chosen for a dramatic reading can therefore produce a voice that sounds right and performs differently.
Choosing an unremarkable clip is counter-intuitive and produces steadier results across a season.
5The cap is per call, not per project
Fifteen seconds in total applies to each generation, so a scene with three speaking characters splits that budget three ways. For a conversation, the practical consequence is that voices are usually supplied one per shot rather than all at once.
That in turn means the shot list has to record which voice belongs to which shot, because the call cannot hold the whole cast.
6Sources read for this entry
This page defines the term and logs what each vendor documents about it, each figure read on 2026-09-12. Related: Dialogue timing, Lip-sync.