Sentioscope

Speech and voice controls, as each vendor documents them

What fifteen seconds of reference audio can carry

Fifteen seconds is the tightest published window for supplying a voice and it appears twice, from two vendors. Elsewhere the figures run to five minutes of sample and twenty minutes of enrolment, which is a different kind of ask entirely. As of 2026-09-12.

What a very short reference clip can and cannot carryFifteen seconds establishes what a voice sounds like and not how it acts. Accent, pacing and emotional register travel unpredictably, and no vendor here documents which of them a model attends to.Transfers reliablyTravels unpredictablyWhat it isTimbre, the colour of the voiceAccent, pacing, emotional registerWhyA short clean sample is enoughThey need a performance, not a sampleWhat to chooseOrdinary speech, cleanly recordedNothing; do not choose for theseDocumented by anyoneNo vendor says soNo vendor says soFifteen seconds appears twice, from two vendors
Fig. 1 An unremarkable clip usually outperforms a dramatic one, because a stable target beats whatever survived a trim.
Published windows for audio supplied so speech follows a chosen voice. Recorded 2026-09-12.
ModelWhat is documentedDetail
HeyGenOne recording, or twenty minutes and upInstant grade, or professional
MiniMaxFifteen seconds in total, at most three clipsPer generation, not per project
RunwayTen seconds to five minutes, up to 10 MBOr a description of twenty characters
Wan 3.0One to fifteen seconds a clip, fifteen in totalWAV or MP3, up to 15 MB

Inclusion rule. Entries whose documentation publishes a duration for audio supplied as the source of a voice. A vendor offering a catalogue without accepting a recording does not earn a row. Order. Alphabetical by model name.

1Timbre transfers, and the rest of a performance may not

Fifteen seconds is comfortably enough to establish what a voice sounds like and not enough to establish how it acts. Accent, pacing and emotional register travel unpredictably, and none of the vendors here documents which of those the model attends to.

The counter-intuitive consequence is that an unremarkable clip usually outperforms a dramatic one. A clean stretch of ordinary speech gives a model a stable target; a theatrical excerpt that had to be trimmed gives it whatever survived the trim.

2A cap per call and a cap per project are different budgets

Where the fifteen seconds applies to each generation, a scene with three speaking characters divides it three ways. In practice that means voices are supplied one per shot, and the shot list becomes the record of which voice belongs where.

Where a voice is enrolled once, the window is spent at the start and never again. Twenty minutes of studio time is a much larger ask and a much smaller ongoing one, which is the trade a series has to price.

3Keep the clip with the character, not with the shot

Whatever audio produced a voice has to live beside that character's visual reference under a name that says so. A shot regenerated months later with a different clip returns a different person, and nothing in the returned file records the substitution.

This is the least glamorous discipline in generated dialogue and the one that decides whether a season holds together. No vendor here documents a way to store the clip for you, so it is an archival job by default.

  • Voice source
    Reference audio is capped at 15 seconds in total across at most 3 clipsa hard published limitMiniMax, video generation guide / recorded 2026-09-12
  • Voice source
    Reference audio is WAV or MP3, one to fifteen seconds a clip, fifteen seconds in total and no more than 15 MBa published capAlibaba Cloud, Wan 3.0 API reference / recorded 2026-09-22
  • Voice source
    A custom voice can be built from an audio sample between 10 seconds and 5 minutes long and at most 10 MBa published sample windowRunway, custom voices reference / recorded 2026-09-22
  • Voice source
    A voice clone is instant from one recording or professional from twenty minutes or more, and is then passed as a voice idtwo enrolment routes with a thresholdHeyGen, API quick start / recorded 2026-09-22
  • Voice source
    A voice can instead be asked for in words, with the description required to run at least 20 charactersa route that needs no recordingRunway, custom voices reference / recorded 2026-09-22

4Sources

Every line is read from the vendor documentation linked on the speech controls page, checked 2026-09-12. What counts as documented is on how read. Nearby: Answers software can check, What silence costs, Where a line goes.