Sentioscope

Speech and voice controls, as each vendor documents them

MiniMax: the mouth belongs to whoever is on screen

Lip movement is documented as following the speaker in frame. That is the only sentence in this register that touches the commonest shot in drama, and it describes a decision the model makes rather than one a prompt can set. As of 2026-09-12.

The model choosing a face, with no override offeredLip movement following the on-screen speaker implies a selection inside the model. Because it is documented as behaviour there is no parameter to reach for when the selection goes the wrong way.How a production controls which mouth movesTwo faces in frameThe model decidesDocumented as behaviour, so no promptoverrides itA single, voice off-frameYou decideThe ambiguity is removed by coveragerather than by a parameterThe one sentence here about the commonest shot in drama
Fig. 1 Framing becomes the only lever, which makes this a shot-list decision that the documentation never spells out.
MiniMax on lip-sync, statement by statement. Read from the vendor's video generation guide on 2026-09-12.
What the documentation settlesWhat it leaves to a take
Native speech is generated with the picture, with lip-sync tied to the speaker on screenHow the model picks between two candidate faces
Reference audio is capped at fifteen seconds in total across at most three clipsWhether a supplied voice influences which face is chosen
A first-frame image and reference images cannot be used in the same callHow a two-hander is continued from a known frame

Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.

1A behaviour is not a control, and the difference is practical

Saying the mouth follows the on-screen speaker implies a selection is happening inside the model. Since it is documented as behaviour, there is no parameter to reach for when the selection is wrong, and no way to say which of two characters should be talking.

Framing becomes the only lever. A single with the other voice off-frame removes the choice entirely, which is a shot-list decision rather than a prompt one, and the documentation does not spell it out.

2The fullest answer in this column is still one sentence

Grading this cell above a bare capability claim is right and should not be mistaken for confidence. The sentence says which face moves, not how well, not on what evidence, and not what happens when two people overlap, which is what a real conversation does.

What it does earn is a testable prediction. A production can shoot the ambiguous case deliberately and find out, and the result is worth writing down because the page will not tell anyone.

3The others that name what the mouth is following

Four more entries name a driver. Three of them are handed a recording, which makes the driver obvious; this is the only one making the sound itself and still naming what the mouth follows.

  • Audio source
    Native speech with lip-sync tied to the speaker who is on screenspeaker-awareMiniMax, video generation guide / recorded 2026-09-12
  • Voice source
    Reference audio is capped at 15 seconds in total across at most 3 clipsa hard published limitMiniMax, video generation guide / recorded 2026-09-12
  • Constraint that collides with voice work
    A first-frame image and reference images cannot be used in the same callMiniMax, video generation guide / recorded 2026-09-12

4Sources

Read from the video generation guide at platform.minimax.io on 2026-09-12. The same column across every entry is on lip-sync; everything this vendor publishes about speech is on MiniMax. What counts as documented is on how read.