MiniMax and Sora 2: two speakers in one frame
The commonest shot in drama goes unaddressed by most of this register. These two touch it: one ties lip movement to whoever is on screen, the other asks for speakers to be labelled with turns alternated. As of 2026-09-22.
| Field | MiniMax | Sora 2 | Where they part |
|---|---|---|---|
| Audio source | With the picture | With the picture | Same answer |
| Voice source | A reference clip on the call | Nothing documented as an input | Different answers |
| Per-character binding | On the call | On the prompt | Different answers |
| Languages | Not documented by the vendor | Not documented by the vendor | Same answer |
| Lip-sync | Whoever is on screen | Named only where it fails | Different answers |
Inclusion rule. Two entries are given a page together when at least one column puts them at opposite grades of answer. Pairs that agree on every column, or that are both blank throughout, do not get a page. Order. Fixed field order, identical on every side-by-side page.
1A behaviour and an instruction are different kinds of help
Tying lip movement to the on-screen speaker says the model makes a choice, with no parameter to override it. Asking for consistent labels and alternating turns hands the choice to the writer, inside the prompt.
The second is more controllable and the first needs less discipline. A production working with the first shapes coverage instead: a single with the other voice off-frame removes the ambiguity entirely.
2Neither publishes a mechanism, and one publishes a limit
Neither says what actually drives a mouth. One adds a warning that long, complex speeches are unlikely to sync, which is an admission about where alignment stops holding and the most actionable sentence either offers.
Read together they point the same way: keep lines short, keep turns explicit, and keep the other speaker out of frame when it matters. None of that is in either document as advice, and both support it.
3Their voice columns are opposite
One takes a reference clip with a published cap and ties a voice to it per call. The other publishes nothing about where a voice comes from, so a label routes a line to a face without establishing an identity that survives the next call.
For a serial that is the deciding difference between them. One lets a production own the voice badly; the other does not let it own the voice at all.
4Each of them on its own
The column this pair was chosen for is lip-sync, and each entry has a page of its own on it. The full row for either, all five columns with the wording behind each cell, is on its model note.
- MiniMax on lip-sync — whoever is on screen.
- Sora 2 on lip-sync — named only where it fails.
- MiniMax, all five fields — read from the video generation guide.
- Sora 2, all five fields — read from the model page and prompting guide.
- Audio sourceNative speech with lip-sync tied to the speaker who is on screenspeaker-aware
- Voice sourceReference audio is capped at 15 seconds in total across at most 3 clipsa hard published limit
- Two characters in frameFor multi-character scenes the guide asks for speakers labelled consistently and turns alternated, so each line lands on the right characteraddressed rather than skipped
- Dialogue timing against clip lengthA four-second shot is said to accommodate one or two short exchanges and an eight-second clip a few more, with long speeches unlikely to sync
- LanguagesNot documented by the vendor (as of 2026-09-22)neither a list nor a count
5Sources
Every cell above is read from the documentation each vendor publishes, on the dates carried by the two model notes. The whole register on one column is on the field note; entries grouped by the shape of their answer are on the routes. Other pairs: Hedra and Synthesia, Wan 3.0 and Wan2.2-S2V.