Sentioscope

Speech and voice controls, as each vendor documents them

MiniMax and Sora 2: two speakers in one frame

The commonest shot in drama goes unaddressed by most of this register. These two touch it: one ties lip movement to whoever is on screen, the other asks for speakers to be labelled with turns alternated. As of 2026-09-22.

A behaviour, and an instruction to the writerTying lip movement to the on-screen speaker says the model chooses, with no parameter to override it. Asking for consistent labels with turns alternated hands the choice to the writer, inside the prompt.Two people in frame, one of them talkingMiniMaxThe model choosesDocumented as behaviour, so coverage isthe only leverSora 2The prompt choosesLabels route the lines, and long speechesare warned againstNeither publishes a mechanism, and one publishes a limit
Fig. 1 Read together they point one way: short lines, explicit turns, and the other speaker out of frame when it matters.
Controls documented, model by modelFilled where the model documents that control, hollow where nothing is published about it.Controls documented, model by modelMiniMaxSora 2Where they partAudio sourceAudio source — MiniMax: With the pictureAudio source — Sora 2: With the pictureAudio source — Where they part: Same answerVoice sourceVoice source — MiniMax: A reference clip on the callVoice source — Sora 2: Nothing documented as an inputVoice source — Where they part: Different answersPer-character bindingPer-character binding — MiniMax: On the callPer-character binding — Sora 2: On the promptPer-character binding — Where they part: Different answersLanguagesLanguages — MiniMax: Not documented by the vendorLanguages — Sora 2: Not documented by the vendorLanguages — Where they part: Same answerLip-syncLip-sync — MiniMax: Whoever is on screenLip-sync — Sora 2: Named only where it failsLip-sync — Where they part: Different answers
Fig. 2 Filled where the model documents that control, hollow where nothing is published about it.
Models documenting each controlHow many models document each control. A hollow column is a statement about documentation, not capability.Models documenting each controlMiniMax4 of 5Sora 23 of 5Where they part5 of 5
Fig. 3 How many models document each control. A hollow column is a statement about documentation, not capability.
MiniMax and Sora 2 on the five fields this register keeps, with the gap on each. Recorded 2026-09-22.
FieldMiniMaxSora 2Where they part
Audio sourceWith the pictureWith the pictureSame answer
Voice sourceA reference clip on the callNothing documented as an inputDifferent answers
Per-character bindingOn the callOn the promptDifferent answers
LanguagesNot documented by the vendorNot documented by the vendorSame answer
Lip-syncWhoever is on screenNamed only where it failsDifferent answers

Inclusion rule. Two entries are given a page together when at least one column puts them at opposite grades of answer. Pairs that agree on every column, or that are both blank throughout, do not get a page. Order. Fixed field order, identical on every side-by-side page.

1A behaviour and an instruction are different kinds of help

Tying lip movement to the on-screen speaker says the model makes a choice, with no parameter to override it. Asking for consistent labels and alternating turns hands the choice to the writer, inside the prompt.

The second is more controllable and the first needs less discipline. A production working with the first shapes coverage instead: a single with the other voice off-frame removes the ambiguity entirely.

2Neither publishes a mechanism, and one publishes a limit

Neither says what actually drives a mouth. One adds a warning that long, complex speeches are unlikely to sync, which is an admission about where alignment stops holding and the most actionable sentence either offers.

Read together they point the same way: keep lines short, keep turns explicit, and keep the other speaker out of frame when it matters. None of that is in either document as advice, and both support it.

3Their voice columns are opposite

One takes a reference clip with a published cap and ties a voice to it per call. The other publishes nothing about where a voice comes from, so a label routes a line to a face without establishing an identity that survives the next call.

For a serial that is the deciding difference between them. One lets a production own the voice badly; the other does not let it own the voice at all.

4Each of them on its own

The column this pair was chosen for is lip-sync, and each entry has a page of its own on it. The full row for either, all five columns with the wording behind each cell, is on its model note.

  • Audio source
    Native speech with lip-sync tied to the speaker who is on screenspeaker-awareMiniMax, video generation guide / recorded 2026-09-12
  • Voice source
    Reference audio is capped at 15 seconds in total across at most 3 clipsa hard published limitMiniMax, video generation guide / recorded 2026-09-12
  • Two characters in frame
    For multi-character scenes the guide asks for speakers labelled consistently and turns alternated, so each line lands on the right characteraddressed rather than skippedOpenAI, Sora 2 prompting guide / recorded 2026-09-22
  • Dialogue timing against clip length
    A four-second shot is said to accommodate one or two short exchanges and an eight-second clip a few more, with long speeches unlikely to syncOpenAI, Sora 2 prompting guide / recorded 2026-09-22
  • Languages
    Not documented by the vendor (as of 2026-09-22)neither a list nor a countOpenAI, video generation guide / recorded 2026-09-22

5Sources

Every cell above is read from the documentation each vendor publishes, on the dates carried by the two model notes. The whole register on one column is on the field note; entries grouped by the shape of their answer are on the routes. Other pairs: Hedra and Synthesia, Wan 3.0 and Wan2.2-S2V.