What fifteen seconds of reference audio can carry
Fifteen seconds is the tightest published window for supplying a voice and it appears twice, from two vendors. Elsewhere the figures run to five minutes of sample and twenty minutes of enrolment, which is a different kind of ask entirely. As of 2026-09-12.
| Model | What is documented | Detail |
|---|---|---|
| HeyGen | One recording, or twenty minutes and up | Instant grade, or professional |
| MiniMax | Fifteen seconds in total, at most three clips | Per generation, not per project |
| Runway | Ten seconds to five minutes, up to 10 MB | Or a description of twenty characters |
| Wan 3.0 | One to fifteen seconds a clip, fifteen in total | WAV or MP3, up to 15 MB |
Inclusion rule. Entries whose documentation publishes a duration for audio supplied as the source of a voice. A vendor offering a catalogue without accepting a recording does not earn a row. Order. Alphabetical by model name.
1Timbre transfers, and the rest of a performance may not
Fifteen seconds is comfortably enough to establish what a voice sounds like and not enough to establish how it acts. Accent, pacing and emotional register travel unpredictably, and none of the vendors here documents which of those the model attends to.
The counter-intuitive consequence is that an unremarkable clip usually outperforms a dramatic one. A clean stretch of ordinary speech gives a model a stable target; a theatrical excerpt that had to be trimmed gives it whatever survived the trim.
2A cap per call and a cap per project are different budgets
Where the fifteen seconds applies to each generation, a scene with three speaking characters divides it three ways. In practice that means voices are supplied one per shot, and the shot list becomes the record of which voice belongs where.
Where a voice is enrolled once, the window is spent at the start and never again. Twenty minutes of studio time is a much larger ask and a much smaller ongoing one, which is the trade a series has to price.
3Keep the clip with the character, not with the shot
Whatever audio produced a voice has to live beside that character's visual reference under a name that says so. A shot regenerated months later with a different clip returns a different person, and nothing in the returned file records the substitution.
This is the least glamorous discipline in generated dialogue and the one that decides whether a season holds together. No vendor here documents a way to store the clip for you, so it is an archival job by default.
- Voice sourceReference audio is capped at 15 seconds in total across at most 3 clipsa hard published limit
- Voice sourceReference audio is WAV or MP3, one to fifteen seconds a clip, fifteen seconds in total and no more than 15 MBa published cap
- Voice sourceA custom voice can be built from an audio sample between 10 seconds and 5 minutes long and at most 10 MBa published sample window
- Voice sourceA voice clone is instant from one recording or professional from twenty minutes or more, and is then passed as a voice idtwo enrolment routes with a threshold
- Voice sourceA voice can instead be asked for in words, with the description required to run at least 20 charactersa route that needs no recording
4Sources
Every line is read from the vendor documentation linked on the speech controls page, checked 2026-09-12. What counts as documented is on how read. Nearby: Answers software can check, What silence costs, Where a line goes.