Runway and MiniMax: a stored voice, or a clip each time
Runway creates a voice asynchronously, brings it to a ready state with a preview, and addresses it afterwards by id. MiniMax establishes a voice from at most fifteen seconds of reference audio, on every call that needs it. As of 2026-09-22.
| Field | Runway | MiniMax | Where they part |
|---|---|---|---|
| Audio source | Not documented by the vendor | With the picture | Only MiniMax answers |
| Voice source | A 10-second to 5-minute sample, or a sentence | A reference clip on the call | Different answers |
| Per-character binding | On a stored voice | On the call | Different answers |
| Languages | Not documented by the vendor | Not documented by the vendor | Same answer |
| Lip-sync | Not documented by the vendor | Whoever is on screen | Only MiniMax answers |
Inclusion rule. Two entries are given a page together when at least one column puts them at opposite grades of answer. Pairs that agree on every column, or that are both blank throughout, do not get a page. Order. Fixed field order, identical on every side-by-side page.
1Auditioning once against reconstructing forty times
A voice that reaches a ready state with a preview can be heard, approved or rejected before a single frame is billed. A voice rebuilt from a clip is never auditioned as itself; each generation is a fresh attempt at the same voice, and the differences between attempts are silent.
Over a season the second arrangement is the one that drifts. Nothing in a returned file records which clip produced the voice in it, so the drift is only visible to an audience comparing episode one with episode nine.
2One takes five minutes of audio, the other fifteen seconds
The published windows are an order of magnitude apart. Ten seconds to five minutes leaves room for range and makes the choice of excerpt a directing decision. Fifteen seconds in total, split across at most three clips, carries timbre and little else.
That difference changes what a production records. A long window rewards a considered session; a fifteen-second cap rewards a clean piece of ordinary speech, because the trim is what the model actually hears.
3Neither publishes a language, and only one publishes an audio source
The voice-construction entry says nothing about where sound comes from in a shot, because its page is about voices rather than performances. The native-speech entry says where sound comes from and ties lip movement to the on-screen speaker.
Between them they cover the two halves a dialogue scene needs and neither covers both, which is the shape most of this register takes.
4Each of them on its own
The column this pair was chosen for is voice source, and each entry has a page of its own on it. The full row for either, all five columns with the wording behind each cell, is on its model note.
- Runway on voice source — a 10-second to 5-minute sample, or a sentence.
- MiniMax on voice source — a reference clip on the call.
- Runway, all five fields — read from the custom voices reference.
- MiniMax, all five fields — read from the video generation guide.
- Voice sourceA custom voice can be built from an audio sample between 10 seconds and 5 minutes long and at most 10 MBa published sample window
- Per-character bindingA voice is created asynchronously, reaches a ready state with a preview, and is then referred to by its own ida stored voice object
- Voice sourceA voice can instead be asked for in words, with the description required to run at least 20 charactersa route that needs no recording
- Voice sourceReference audio is capped at 15 seconds in total across at most 3 clipsa hard published limit
- Audio sourceNative speech with lip-sync tied to the speaker who is on screenspeaker-aware
5Sources
Every cell above is read from the documentation each vendor publishes, on the dates carried by the two model notes. The whole register on one column is on the field note; entries grouped by the shape of their answer are on the routes. Other pairs: HeyGen and LTX Studio, MiniMax and Wan 3.0.