Sentioscope

Speech and voice controls, as each vendor documents them

Runway and MiniMax: a stored voice, or a clip each time

Runway creates a voice asynchronously, brings it to a ready state with a preview, and addresses it afterwards by id. MiniMax establishes a voice from at most fifteen seconds of reference audio, on every call that needs it. As of 2026-09-22.

Auditioning once, against rebuilding forty timesA voice that reaches a ready state with a preview can be approved before a frame is billed. A voice rebuilt from a clip is never auditioned as itself, and the differences between attempts are silent.RunwayOne sample of 10 s to5 min, previewed,then reused by id.Approved onceMiniMaxFifteen seconds intotal, supplied againon every call.Rebuilt eachtimeOver a seasonOne of the two candrift without anysingle shot lookingwrong.
Fig. 1 The published sample windows are an order of magnitude apart, which changes what a production should record.
Controls documented, model by modelFilled where the model documents that control, hollow where nothing is published about it.Controls documented, model by modelRunwayMiniMaxWhere they partAudio sourceAudio source — Runway: Not documented by the vendorAudio source — MiniMax: With the pictureAudio source — Where they part: Only MiniMax answersVoice sourceVoice source — Runway: A 10-second to 5-minute sample, or a sentenceVoice source — MiniMax: A reference clip on the callVoice source — Where they part: Different answersPer-character bindingPer-character binding — Runway: On a stored voicePer-character binding — MiniMax: On the callPer-character binding — Where they part: Different answersLanguagesLanguages — Runway: Not documented by the vendorLanguages — MiniMax: Not documented by the vendorLanguages — Where they part: Same answerLip-syncLip-sync — Runway: Not documented by the vendorLip-sync — MiniMax: Whoever is on screenLip-sync — Where they part: Only MiniMax answers
Fig. 2 Filled where the model documents that control, hollow where nothing is published about it.
Models documenting each controlHow many models document each control. A hollow column is a statement about documentation, not capability.Models documenting each controlRunway2 of 5MiniMax4 of 5Where they part5 of 5
Fig. 3 How many models document each control. A hollow column is a statement about documentation, not capability.
Runway and MiniMax on the five fields this register keeps, with the gap on each. Recorded 2026-09-22.
FieldRunwayMiniMaxWhere they part
Audio sourceNot documented by the vendorWith the pictureOnly MiniMax answers
Voice sourceA 10-second to 5-minute sample, or a sentenceA reference clip on the callDifferent answers
Per-character bindingOn a stored voiceOn the callDifferent answers
LanguagesNot documented by the vendorNot documented by the vendorSame answer
Lip-syncNot documented by the vendorWhoever is on screenOnly MiniMax answers

Inclusion rule. Two entries are given a page together when at least one column puts them at opposite grades of answer. Pairs that agree on every column, or that are both blank throughout, do not get a page. Order. Fixed field order, identical on every side-by-side page.

1Auditioning once against reconstructing forty times

A voice that reaches a ready state with a preview can be heard, approved or rejected before a single frame is billed. A voice rebuilt from a clip is never auditioned as itself; each generation is a fresh attempt at the same voice, and the differences between attempts are silent.

Over a season the second arrangement is the one that drifts. Nothing in a returned file records which clip produced the voice in it, so the drift is only visible to an audience comparing episode one with episode nine.

2One takes five minutes of audio, the other fifteen seconds

The published windows are an order of magnitude apart. Ten seconds to five minutes leaves room for range and makes the choice of excerpt a directing decision. Fifteen seconds in total, split across at most three clips, carries timbre and little else.

That difference changes what a production records. A long window rewards a considered session; a fifteen-second cap rewards a clean piece of ordinary speech, because the trim is what the model actually hears.

3Neither publishes a language, and only one publishes an audio source

The voice-construction entry says nothing about where sound comes from in a shot, because its page is about voices rather than performances. The native-speech entry says where sound comes from and ties lip movement to the on-screen speaker.

Between them they cover the two halves a dialogue scene needs and neither covers both, which is the shape most of this register takes.

4Each of them on its own

The column this pair was chosen for is voice source, and each entry has a page of its own on it. The full row for either, all five columns with the wording behind each cell, is on its model note.

  • Voice source
    A custom voice can be built from an audio sample between 10 seconds and 5 minutes long and at most 10 MBa published sample windowRunway, custom voices reference / recorded 2026-09-22
  • Per-character binding
    A voice is created asynchronously, reaches a ready state with a preview, and is then referred to by its own ida stored voice objectRunway, custom voices reference / recorded 2026-09-22
  • Voice source
    A voice can instead be asked for in words, with the description required to run at least 20 charactersa route that needs no recordingRunway, custom voices reference / recorded 2026-09-22
  • Voice source
    Reference audio is capped at 15 seconds in total across at most 3 clipsa hard published limitMiniMax, video generation guide / recorded 2026-09-12
  • Audio source
    Native speech with lip-sync tied to the speaker who is on screenspeaker-awareMiniMax, video generation guide / recorded 2026-09-12

5Sources

Every cell above is read from the documentation each vendor publishes, on the dates carried by the two model notes. The whole register on one column is on the field note; entries grouped by the shape of their answer are on the routes. Other pairs: HeyGen and LTX Studio, MiniMax and Wan 3.0.