MiniMax: native speech that picks a speaker
MiniMax documents native speech arriving with the picture and ties lip movement to the speaker on screen. That second half is a behaviour rather than a control: the model is deciding which face moves, and no parameter overrides it. As of 2026-09-12.
| What the documentation settles | What it leaves to a take |
|---|---|
| Native speech is generated with the picture, with lip-sync tied to the speaker on screen | How the model chooses between two candidate faces in one frame |
| Reference audio is capped at fifteen seconds in total across at most three clips | Whether fifteen seconds is enough to fix an accent as well as a timbre |
| A first-frame image and reference images cannot be used in the same call | How a conversation between established characters is continued from a known frame |
Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.
1Speaker-aware alignment is a promise about the hard case
Most vendors describe audio as an output and let the reader assume the mouths work out. This one says lip movement follows whoever is in frame, which is a claim about the shot that breaks most often: two people, one of them talking. It is documented as a behaviour, so there is no parameter to reach for when it picks wrongly.
That makes the framing decision load-bearing. A two-shot hands the model a choice; a single with the other voice off-frame does not. The documentation does not say that, but it is the practical consequence of the sentence it does publish.
2A constraint from elsewhere lands on dialogue work
The rule that a first-frame image and reference images cannot travel together is not about audio at all, and it collides with dialogue anyway. Continuing from a known frame and carrying a character reference are exactly the two things a conversation between established characters wants at the same time.
It is recorded in this field because a register that listed only the audio sentences would describe a workflow the documentation forbids. The cell answers where the sound comes from; the constraint decides what else can be in the call that produces it.
3The others that make sound while they make the picture
Seven more entries generate audio in the same pass. This is the only one whose documentation says anything about which face the sound belongs to.
- Kling AI — with the picture, on video 3.0.
- LTX Studio — with the picture, and audio to video.
- SceneMixer — with the picture, in the language set for the project.
- Sora 2 — with the picture.
- Veo — with the picture.
- Vidu — with the picture, speech available on its own.
- Wan 3.0 — with the picture, unless switched off.
- Audio sourceNative speech with lip-sync tied to the speaker who is on screenspeaker-aware
- Voice sourceReference audio is capped at 15 seconds in total across at most 3 clipsa hard published limit
- Constraint that collides with voice workA first-frame image and reference images cannot be used in the same call
4Sources
Read from the video generation guide at platform.minimax.io on 2026-09-12. The same column across every entry is on audio source; everything this vendor publishes about speech is on MiniMax. What counts as documented is on how read.