Sentioscope

Speech and voice controls, as each vendor documents them

MiniMax: the clip goes in again, every single time

A voice is established from a reference clip on the call that uses it, within fifteen seconds across at most three clips. There is no documented way to store the result, so every generation is an opportunity for the voice to move. As of 2026-09-12.

Rebuilding a voice forty times, quietly differentlyA stored voice is wrong once. A voice rebuilt from a clip on every call can move each time, and nothing in the returned file records which clip produced it, so a season drifts without any single shot looking wrong.Call oneFifteen seconds ofreference audioestablishes thevoice.Clip discardedCall twentyThe same fifteenseconds, or whicheverfile was at hand.Clip discardedCall fortyNothing on theplatform rememberswhat the first twoused.
Fig. 1 The defence is archival rather than technical: keep the clip, name it after the character, and never reach for whatever is convenient.
MiniMax on per-character binding, statement by statement. Read from the vendor's video generation guide on 2026-09-12.
What the documentation settlesWhat it leaves to a take
Reference audio is capped at fifteen seconds in total across at most three clipsHow much of a voice survives being rebuilt from the same clip twice
A reference clip, capped across three clips, is supplied each timeWhether an identical clip produces an identical voice
A first-frame image and reference images cannot be used in the same callHow a continuing conversation supplies its voices at all

Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.

1Per-call reconstruction has a silent failure mode

A stored voice is wrong once. A rebuilt voice can be slightly different on every generation, and the difference is silent: nothing in the returned file records which clip produced it, so a season can drift without any single shot looking wrong.

The defence is archival rather than technical. The clip that produced a character has to be kept, named and used unchanged, and a regenerated shot months later has to reach for the same file rather than for whatever is convenient.

2The exclusivity rule collides with a conversation

If a first-frame image and reference images cannot be used together, then continuing from a known frame and carrying character references are two things one call cannot do. A dialogue scene between established characters wants both at once.

Recorded in this column because it decides how a cast is actually supplied. In practice voices go in one per shot, and the shot list has to say which voice belongs to which shot, which is bookkeeping that a bound voice would have removed.

3The others that re-establish the voice on each call

Four more entries put the voice on the request. Three of them pass a short stable identifier; this one and one other pass audio, which is the version that can drift.

  • Voice source
    Reference audio is capped at 15 seconds in total across at most 3 clipsa hard published limitMiniMax, video generation guide / recorded 2026-09-12
  • Constraint that collides with voice work
    A first-frame image and reference images cannot be used in the same callMiniMax, video generation guide / recorded 2026-09-12
  • Audio source
    Native speech with lip-sync tied to the speaker who is on screenspeaker-awareMiniMax, video generation guide / recorded 2026-09-12

4Sources

Read from the video generation guide at platform.minimax.io on 2026-09-12. The same column across every entry is on per-character binding; everything this vendor publishes about speech is on MiniMax. What counts as documented is on how read.