Sentioscope

Speech and voice controls, as each vendor documents them

MiniMax and Wan 3.0: the same fifteen seconds, differently

Both cap reference audio at fifteen seconds in total. One divides that across at most three clips; the other lets a single clip fill it and publishes the format and the byte ceiling beside it. As of 2026-09-22.

The same fifteen seconds, arranged two waysOne divides the total across at most three clips; the other lets a single clip fill it and publishes the format and the byte ceiling beside it. Three samples and one considered take are different asks.MiniMaxWan 3.0The totalFifteen secondsFifteen secondsHow it is splitAcross at most three clipsOne clip may use all of itFormat and sizeWAV or MP3WAV or MP3, up to 15 MBNearby constraintFirst frame and references cannot share acallAudio off costs the same as onIdentical caps, and two different workflows
Fig. 1 A number that recurs across vendors usually reflects a shared constraint rather than a shared decision.
Controls documented, model by modelFilled where the model documents that control, hollow where nothing is published about it.Controls documented, model by modelMiniMaxWan 3.0Where they partAudio sourceAudio source — MiniMax: With the pictureAudio source — Wan 3.0: With the picture, unless switched offAudio source — Where they part: Different answersVoice sourceVoice source — MiniMax: A reference clip on the callVoice source — Wan 3.0: Reference audio, 15 seconds in totalVoice source — Where they part: Different answersPer-character bindingPer-character binding — MiniMax: On the callPer-character binding — Wan 3.0: On the callPer-character binding — Where they part: Same answerLanguagesLanguages — MiniMax: Not documented by the vendorLanguages — Wan 3.0: Not documented by the vendorLanguages — Where they part: Same answerLip-syncLip-sync — MiniMax: Whoever is on screenLip-sync — Wan 3.0: Not documented by the vendorLip-sync — Where they part: Only MiniMax answers
Fig. 2 Filled where the model documents that control, hollow where nothing is published about it.
Models documenting each controlHow many models document each control. A hollow column is a statement about documentation, not capability.Models documenting each controlMiniMax4 of 5Wan 3.03 of 5Where they part5 of 5
Fig. 3 How many models document each control. A hollow column is a statement about documentation, not capability.
MiniMax and Wan 3.0 on the five fields this register keeps, with the gap on each. Recorded 2026-09-22.
FieldMiniMaxWan 3.0Where they part
Audio sourceWith the pictureWith the picture, unless switched offDifferent answers
Voice sourceA reference clip on the callReference audio, 15 seconds in totalDifferent answers
Per-character bindingOn the callOn the callSame answer
LanguagesNot documented by the vendorNot documented by the vendorSame answer
Lip-syncWhoever is on screenNot documented by the vendorOnly MiniMax answers

Inclusion rule. Two entries are given a page together when at least one column puts them at opposite grades of answer. Pairs that agree on every column, or that are both blank throughout, do not get a page. Order. Fixed field order, identical on every side-by-side page.

1A recurring number usually means a shared constraint

Two vendors landing on the same fifteen seconds, independently, suggests the figure reflects something about how much audio a model needs rather than a product decision. For a production that is useful: fifteen seconds can be treated as the planning unit for reference audio generally.

Choose a clean fifteen seconds of ordinary speech per character, store it beside the visual reference, and it will fit wherever a clip is accepted. The arrangement around the cap differs; the cap does not.

2One clean take against three short ones

Dividing fifteen seconds across three clips and allowing one clip to use all fifteen are different asks. Three clips invite a range of samples and make the trim decisions three times; one long clip rewards a single considered take.

Neither vendor says which produces a better voice, and it would be a reasonable thing to test once per production rather than per character. The result is worth writing down, because nothing in either reference predicts it.

3The surrounding columns diverge sharply

One ties lip movement to the on-screen speaker, which is the only sentence in this register about the commonest shot in drama. The other says nothing about mouths and does say that enabling or disabling audio does not affect pricing.

One also forbids a first-frame image and reference images in the same call, which collides with continuing a conversation. Identical caps, and two quite different workflows.

4Each of them on its own

The column this pair was chosen for is voice source, and each entry has a page of its own on it. The full row for either, all five columns with the wording behind each cell, is on its model note.

5Sources

Every cell above is read from the documentation each vendor publishes, on the dates carried by the two model notes. The whole register on one column is on the field note; entries grouped by the shape of their answer are on the routes. Other pairs: Sora 2 and Luma Ray, Synthesia and D-ID.