MiniMax: the clip goes in again, every single time
A voice is established from a reference clip on the call that uses it, within fifteen seconds across at most three clips. There is no documented way to store the result, so every generation is an opportunity for the voice to move. As of 2026-09-12.
| What the documentation settles | What it leaves to a take |
|---|---|
| Reference audio is capped at fifteen seconds in total across at most three clips | How much of a voice survives being rebuilt from the same clip twice |
| A reference clip, capped across three clips, is supplied each time | Whether an identical clip produces an identical voice |
| A first-frame image and reference images cannot be used in the same call | How a continuing conversation supplies its voices at all |
Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.
1Per-call reconstruction has a silent failure mode
A stored voice is wrong once. A rebuilt voice can be slightly different on every generation, and the difference is silent: nothing in the returned file records which clip produced it, so a season can drift without any single shot looking wrong.
The defence is archival rather than technical. The clip that produced a character has to be kept, named and used unchanged, and a regenerated shot months later has to reach for the same file rather than for whatever is convenient.
2The exclusivity rule collides with a conversation
If a first-frame image and reference images cannot be used together, then continuing from a known frame and carrying character references are two things one call cannot do. A dialogue scene between established characters wants both at once.
Recorded in this column because it decides how a cast is actually supplied. In practice voices go in one per shot, and the shot list has to say which voice belongs to which shot, which is bookkeeping that a bound voice would have removed.
3The others that re-establish the voice on each call
Four more entries put the voice on the request. Three of them pass a short stable identifier; this one and one other pass audio, which is the version that can drift.
- Voice sourceReference audio is capped at 15 seconds in total across at most 3 clipsa hard published limit
- Constraint that collides with voice workA first-frame image and reference images cannot be used in the same call
- Audio sourceNative speech with lip-sync tied to the speaker who is on screenspeaker-aware
4Sources
Read from the video generation guide at platform.minimax.io on 2026-09-12. The same column across every entry is on per-character binding; everything this vendor publishes about speech is on MiniMax. What counts as documented is on how read.