MiniMax: fifteen seconds, across at most three clips
Reference audio is capped at fifteen seconds in total, across at most three clips, in WAV or MP3. It is the tightest published limit here and the only one stated as a figure rather than as a capability. As of 2026-09-12.
| What the documentation settles | What it leaves to a take |
|---|---|
| Reference audio is capped at fifteen seconds in total across at most three clips | Whether fifteen seconds fixes an accent as well as a timbre |
| Native speech is generated with the picture, with lip-sync tied to the speaker on screen | Whether a supplied voice changes how that speaker is chosen |
| A first-frame image and reference images cannot be used in the same call | How a voice is supplied while continuing from a known frame |
Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.
1Fifteen seconds is enough for a voice and not for a performance
A cap this tight decides how the clip should be chosen. Fifteen seconds carries timbre reliably and carries pacing, accent and emotional register unpredictably, so a clean piece of ordinary speech outperforms a dramatic excerpt that had to be trimmed. The trim is what the model hears.
It also means the clip is an asset, not a step. Whatever fifteen seconds produced a character has to be stored beside that character, because a different clip months later produces a different person and nothing in the output announces it.
2Per call, not per project, which splits the budget by speaker
Fifteen seconds applies to each generation. A scene with three speaking characters therefore divides that allowance three ways, which in practice means voices are supplied one per shot rather than all at once, and the shot list has to record which voice belongs to which shot.
The exclusivity rule makes the arithmetic worse. If a first-frame image and reference images cannot travel together, then continuing a conversation from a known frame and carrying the voices for it are two things the same call cannot do.
3The others that take a clip on the call that uses it
Two more entries establish the voice from audio supplied with each generation. One publishes an identical fifteen-second total; the other takes the whole performance as its input.
- Wan 3.0 — reference audio, 15 seconds in total.
- Wan2.2-S2V — the audio track itself.
- Voice sourceReference audio is capped at 15 seconds in total across at most 3 clipsa hard published limit
- Audio sourceNative speech with lip-sync tied to the speaker who is on screenspeaker-aware
- Constraint that collides with voice workA first-frame image and reference images cannot be used in the same call
4Sources
Read from the video generation guide at platform.minimax.io on 2026-09-12. The same column across every entry is on voice source; everything this vendor publishes about speech is on MiniMax. What counts as documented is on how read.