MiniMax and Wan 3.0: the same fifteen seconds, differently
Both cap reference audio at fifteen seconds in total. One divides that across at most three clips; the other lets a single clip fill it and publishes the format and the byte ceiling beside it. As of 2026-09-22.
| Field | MiniMax | Wan 3.0 | Where they part |
|---|---|---|---|
| Audio source | With the picture | With the picture, unless switched off | Different answers |
| Voice source | A reference clip on the call | Reference audio, 15 seconds in total | Different answers |
| Per-character binding | On the call | On the call | Same answer |
| Languages | Not documented by the vendor | Not documented by the vendor | Same answer |
| Lip-sync | Whoever is on screen | Not documented by the vendor | Only MiniMax answers |
Inclusion rule. Two entries are given a page together when at least one column puts them at opposite grades of answer. Pairs that agree on every column, or that are both blank throughout, do not get a page. Order. Fixed field order, identical on every side-by-side page.
1A recurring number usually means a shared constraint
Two vendors landing on the same fifteen seconds, independently, suggests the figure reflects something about how much audio a model needs rather than a product decision. For a production that is useful: fifteen seconds can be treated as the planning unit for reference audio generally.
Choose a clean fifteen seconds of ordinary speech per character, store it beside the visual reference, and it will fit wherever a clip is accepted. The arrangement around the cap differs; the cap does not.
2One clean take against three short ones
Dividing fifteen seconds across three clips and allowing one clip to use all fifteen are different asks. Three clips invite a range of samples and make the trim decisions three times; one long clip rewards a single considered take.
Neither vendor says which produces a better voice, and it would be a reasonable thing to test once per production rather than per character. The result is worth writing down, because nothing in either reference predicts it.
3The surrounding columns diverge sharply
One ties lip movement to the on-screen speaker, which is the only sentence in this register about the commonest shot in drama. The other says nothing about mouths and does say that enabling or disabling audio does not affect pricing.
One also forbids a first-frame image and reference images in the same call, which collides with continuing a conversation. Identical caps, and two quite different workflows.
4Each of them on its own
The column this pair was chosen for is voice source, and each entry has a page of its own on it. The full row for either, all five columns with the wording behind each cell, is on its model note.
- MiniMax on voice source — a reference clip on the call.
- Wan 3.0 on voice source — reference audio, 15 seconds in total.
- MiniMax, all five fields — read from the video generation guide.
- Wan 3.0, all five fields — read from the video generation API reference.
- Voice sourceReference audio is capped at 15 seconds in total across at most 3 clipsa hard published limit
- Constraint that collides with voice workA first-frame image and reference images cannot be used in the same call
- Voice sourceReference audio is WAV or MP3, one to fifteen seconds a clip, fifteen seconds in total and no more than 15 MBa published cap
- What choosing silence costsEnabling or disabling audio does not affect pricing
- How a spoken line is written into a promptThe worked example puts the spoken line in the prompt itself, after the word saying
5Sources
Every cell above is read from the documentation each vendor publishes, on the dates carried by the two model notes. The whole register on one column is on the field note; entries grouped by the shape of their answer are on the routes. Other pairs: Sora 2 and Luma Ray, Synthesia and D-ID.