Sentioscope

Speech and voice controls, as each vendor documents them

MiniMax: native speech that picks a speaker

MiniMax documents native speech arriving with the picture and ties lip movement to the speaker on screen. That second half is a behaviour rather than a control: the model is deciding which face moves, and no parameter overrides it. As of 2026-09-12.

Speaker-aware alignment, and the shot that tests itLip movement is documented as following whoever is on screen, which means the model is choosing a face. With one person in frame the choice is trivial; with two it is the decision the shot depends on.Who is in frame when the line landsOne personNo choice to makeAlignment is a formality; the sentencepromises nothing extraTwo peopleThe model picksNo parameter overrides it, so coveragecarries the riskA single with the other voice off-frame removes the choice
Fig. 1 The documentation states the behaviour rather than offering a parameter, so framing is the only lever a production actually holds.
MiniMax on audio source, statement by statement. Read from the vendor's video generation guide on 2026-09-12.
What the documentation settlesWhat it leaves to a take
Native speech is generated with the picture, with lip-sync tied to the speaker on screenHow the model chooses between two candidate faces in one frame
Reference audio is capped at fifteen seconds in total across at most three clipsWhether fifteen seconds is enough to fix an accent as well as a timbre
A first-frame image and reference images cannot be used in the same callHow a conversation between established characters is continued from a known frame

Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.

1Speaker-aware alignment is a promise about the hard case

Most vendors describe audio as an output and let the reader assume the mouths work out. This one says lip movement follows whoever is in frame, which is a claim about the shot that breaks most often: two people, one of them talking. It is documented as a behaviour, so there is no parameter to reach for when it picks wrongly.

That makes the framing decision load-bearing. A two-shot hands the model a choice; a single with the other voice off-frame does not. The documentation does not say that, but it is the practical consequence of the sentence it does publish.

2A constraint from elsewhere lands on dialogue work

The rule that a first-frame image and reference images cannot travel together is not about audio at all, and it collides with dialogue anyway. Continuing from a known frame and carrying a character reference are exactly the two things a conversation between established characters wants at the same time.

It is recorded in this field because a register that listed only the audio sentences would describe a workflow the documentation forbids. The cell answers where the sound comes from; the constraint decides what else can be in the call that produces it.

3The others that make sound while they make the picture

Seven more entries generate audio in the same pass. This is the only one whose documentation says anything about which face the sound belongs to.

  • Kling AI — with the picture, on video 3.0.
  • LTX Studio — with the picture, and audio to video.
  • SceneMixer — with the picture, in the language set for the project.
  • Sora 2 — with the picture.
  • Veo — with the picture.
  • Vidu — with the picture, speech available on its own.
  • Wan 3.0 — with the picture, unless switched off.
  • Audio source
    Native speech with lip-sync tied to the speaker who is on screenspeaker-awareMiniMax, video generation guide / recorded 2026-09-12
  • Voice source
    Reference audio is capped at 15 seconds in total across at most 3 clipsa hard published limitMiniMax, video generation guide / recorded 2026-09-12
  • Constraint that collides with voice work
    A first-frame image and reference images cannot be used in the same callMiniMax, video generation guide / recorded 2026-09-12

4Sources

Read from the video generation guide at platform.minimax.io on 2026-09-12. The same column across every entry is on audio source; everything this vendor publishes about speech is on MiniMax. What counts as documented is on how read.