MiniMax: the mouth belongs to whoever is on screen
Lip movement is documented as following the speaker in frame. That is the only sentence in this register that touches the commonest shot in drama, and it describes a decision the model makes rather than one a prompt can set. As of 2026-09-12.
| What the documentation settles | What it leaves to a take |
|---|---|
| Native speech is generated with the picture, with lip-sync tied to the speaker on screen | How the model picks between two candidate faces |
| Reference audio is capped at fifteen seconds in total across at most three clips | Whether a supplied voice influences which face is chosen |
| A first-frame image and reference images cannot be used in the same call | How a two-hander is continued from a known frame |
Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.
1A behaviour is not a control, and the difference is practical
Saying the mouth follows the on-screen speaker implies a selection is happening inside the model. Since it is documented as behaviour, there is no parameter to reach for when the selection is wrong, and no way to say which of two characters should be talking.
Framing becomes the only lever. A single with the other voice off-frame removes the choice entirely, which is a shot-list decision rather than a prompt one, and the documentation does not spell it out.
2The fullest answer in this column is still one sentence
Grading this cell above a bare capability claim is right and should not be mistaken for confidence. The sentence says which face moves, not how well, not on what evidence, and not what happens when two people overlap, which is what a real conversation does.
What it does earn is a testable prediction. A production can shoot the ambiguous case deliberately and find out, and the result is worth writing down because the page will not tell anyone.
3The others that name what the mouth is following
Four more entries name a driver. Three of them are handed a recording, which makes the driver obvious; this is the only one making the sound itself and still naming what the mouth follows.
- Hedra — the supplied audio.
- sync-3 — the audio it is given.
- Synthesia — the spoken content, framed close.
- Wan2.2-S2V — the audio input.
- Audio sourceNative speech with lip-sync tied to the speaker who is on screenspeaker-aware
- Voice sourceReference audio is capped at 15 seconds in total across at most 3 clipsa hard published limit
- Constraint that collides with voice workA first-frame image and reference images cannot be used in the same call
4Sources
Read from the video generation guide at platform.minimax.io on 2026-09-12. The same column across every entry is on lip-sync; everything this vendor publishes about speech is on MiniMax. What counts as documented is on how read.