Sentioscope

Speech and voice controls, as each vendor documents them

MiniMax: fifteen seconds, across at most three clips

Reference audio is capped at fifteen seconds in total, across at most three clips, in WAV or MP3. It is the tightest published limit here and the only one stated as a figure rather than as a capability. As of 2026-09-12.

Fifteen seconds, and why it is chosen for timbreFifteen seconds in total across at most three clips carries timbre reliably and carries pacing, accent and emotional register unpredictably. The trim is what the model hears, so the excerpt decides the character.Choose the clipClean ordinary speechbeats a dramaticexcerpt that had tobe cut.Up to 15 sSupply it per callThe cap applies toeach generation, notto the project.Split by speakerStore it by nameNothing in thereturned file recordswhich sample spoke.
Fig. 1 Whatever fifteen seconds produced a character has to be stored beside it, because a different clip months later makes a different person.
MiniMax on voice source, statement by statement. Read from the vendor's video generation guide on 2026-09-12.
What the documentation settlesWhat it leaves to a take
Reference audio is capped at fifteen seconds in total across at most three clipsWhether fifteen seconds fixes an accent as well as a timbre
Native speech is generated with the picture, with lip-sync tied to the speaker on screenWhether a supplied voice changes how that speaker is chosen
A first-frame image and reference images cannot be used in the same callHow a voice is supplied while continuing from a known frame

Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.

1Fifteen seconds is enough for a voice and not for a performance

A cap this tight decides how the clip should be chosen. Fifteen seconds carries timbre reliably and carries pacing, accent and emotional register unpredictably, so a clean piece of ordinary speech outperforms a dramatic excerpt that had to be trimmed. The trim is what the model hears.

It also means the clip is an asset, not a step. Whatever fifteen seconds produced a character has to be stored beside that character, because a different clip months later produces a different person and nothing in the output announces it.

2Per call, not per project, which splits the budget by speaker

Fifteen seconds applies to each generation. A scene with three speaking characters therefore divides that allowance three ways, which in practice means voices are supplied one per shot rather than all at once, and the shot list has to record which voice belongs to which shot.

The exclusivity rule makes the arithmetic worse. If a first-frame image and reference images cannot travel together, then continuing a conversation from a known frame and carrying the voices for it are two things the same call cannot do.

3The others that take a clip on the call that uses it

Two more entries establish the voice from audio supplied with each generation. One publishes an identical fifteen-second total; the other takes the whole performance as its input.

  • Wan 3.0 — reference audio, 15 seconds in total.
  • Wan2.2-S2V — the audio track itself.
  • Voice source
    Reference audio is capped at 15 seconds in total across at most 3 clipsa hard published limitMiniMax, video generation guide / recorded 2026-09-12
  • Audio source
    Native speech with lip-sync tied to the speaker who is on screenspeaker-awareMiniMax, video generation guide / recorded 2026-09-12
  • Constraint that collides with voice work
    A first-frame image and reference images cannot be used in the same callMiniMax, video generation guide / recorded 2026-09-12

4Sources

Read from the video generation guide at platform.minimax.io on 2026-09-12. The same column across every entry is on voice source; everything this vendor publishes about speech is on MiniMax. What counts as documented is on how read.