Sentioscope

Speech and voice controls, as each vendor documents them

Entries that publish a limit as a number

Nine entries publish at least one limit as a number rather than as a capability. Seconds, minutes, megabytes and characters. These are the figures a schedule can be built from, and they are the rarest thing in this register. As of 2026-09-22.

Vendors bound their inputs more readily than their outputsLined up in one table the pattern is clear: the tightest published figures are all about what a production hands over, and the loosest are about what comes back. Nine entries publish a number at all.Figures about inputsFigures about outputsTypical sizeSeconds, and tens of megabytesMinutes, and whole generationsHow tightFifteen seconds is commonTen minutes, or a hundred and twentysecondsWhat they boundReference clips, samples, scriptsA clip, a talk, an allowanceWho they constrainA production's materialA vendor's promiseFifteen seconds recurs across vendors, which suggests a shared constraint
Fig. 1 A vendor that says fifteen seconds has said something that can be wrong, which is more than a claim of full support manages.
Published numerical limits across the register, and what each one constrains. Recorded 2026-09-22.
ModelThe published figureWhat it limits
D-IDFive minutes for a clip, ten for a talk; 40,000 characters of scriptSupplied audio, and the length of a text script
HedraA maximum duration of ten minutesThe whole generated video
HeyGenTwenty minutes or more for the professional gradeEnrolling a voice clone
MiniMaxFifteen seconds in total, across at most three clipsReference audio per call
PixVerseSixty seconds and one hundred megabytesAudio and video on the request
RunwayTen seconds to five minutes, at most 10 MBA sample for a custom voice
Sora 2Sixteen or twenty seconds, to a total of 120One generation, and its extensions
sync-3One generation a month, fifteen secondsThe free allowance
Wan 3.0One to fifteen seconds a clip, fifteen in total, 15 MBReference audio per call

Inclusion rule. Entries whose documentation states at least one limit as a figure bearing on speech: a duration, a file size, a character count or an allowance. A capability described without a number does not earn a row. Order. Alphabetical by model name.

1A figure is the part of a page a schedule can actually use

Everything else in this register is prose about capability. A figure is different: it can be divided into an episode count, multiplied by a cast size, or compared against a delivery slot. Nine entries publish one, and the other eight leave every quantity to discovery.

Publishing a figure is also a commitment. A vendor that says fifteen seconds has said something that can be wrong, which is more than a vendor claiming full support has done.

2The figures are not measuring the same thing

Ten minutes of output duration, twenty minutes of enrolment audio and fifteen seconds of reference clip look comparable and are not. One bounds a deliverable, one bounds a studio booking, and one bounds how much evidence a model gets about a voice.

Lined up in one table the pattern is clearer: the tightest figures are all about what a production hands over, and the loosest are all about what comes back. Vendors are more willing to bound their inputs than their outputs.

3Two figures repeat, and the repetition is informative

Fifteen seconds in total appears twice, from two different vendors, with different arrangements around it. Sixty seconds and a hundred megabytes appear together on one request. A number that recurs across vendors usually reflects a shared constraint rather than a shared decision.

For a production the useful consequence is that fifteen seconds can be treated as the planning unit for reference audio generally. Choose a clean fifteen seconds of ordinary speech per character, store it beside the visual reference, and it will fit wherever a clip is accepted.

4The entries on this route, one page each

Each of these links to that entry on the field this route groups by. The wording behind every cell, and the date it was read, sits on the page it links to.

  • D-ID — a voice id, or a recording by url.
  • Hedra — a track, or a voice id from the endpoint.
  • HeyGen — a recording to clone from.
  • MiniMax — a reference clip on the call.
  • PixVerse — a sample, or a built-in voice.
  • Runway — a 10-second to 5-minute sample, or a sentence.
  • Sora 2 — nothing documented as an input.
  • sync-3 — nothing documented as an input.
  • Wan 3.0 — reference audio, 15 seconds in total.
  • Voice source
    Reference audio is capped at 15 seconds in total across at most 3 clipsa hard published limitMiniMax, video generation guide / recorded 2026-09-12
  • Voice source
    Reference audio is WAV or MP3, one to fifteen seconds a clip, fifteen seconds in total and no more than 15 MBa published capAlibaba Cloud, Wan 3.0 API reference / recorded 2026-09-22
  • Voice source
    A custom voice can be built from an audio sample between 10 seconds and 5 minutes long and at most 10 MBa published sample windowRunway, custom voices reference / recorded 2026-09-22
  • Length per call
    Audio and video are each capped at sixty seconds and one hundred megabytesPixVerse, speech and lip sync guide / recorded 2026-09-22
  • Shot length a line has to fit
    Generations run 16 or 20 seconds, and an extension adds up to 20 seconds at a time to a total of 120OpenAI, video generation guide / recorded 2026-09-22
  • Length per call on the free tier
    A free account runs one generation a month with a fifteen-second ceiling, with paid limits following the planSync, sync-3 model documentation / recorded 2026-09-22
  • Audio source
    A script is either text of up to forty thousand characters or an audio url, with audio limited to five minutes for clips and ten for talksa script read aloud or a fileD-ID, create a talk reference / recorded 2026-09-22
  • Shot length set by the take
    The audio normally determines the video length, and omitting the duration follows the source audioHedra, avatar video guide / recorded 2026-09-22
  • Voice source
    A voice clone is instant from one recording or professional from twenty minutes or more, and is then passed as a voice idtwo enrolment routes with a thresholdHeyGen, API quick start / recorded 2026-09-22

5Sources

Membership of this route is decided by the wording in the field it groups by, read from the documentation each vendor publishes. The column itself is on the field note, all five columns are on the speech table, and what counts as documented is on how read. Other routes: A named driver, Named, not explained.