What MiniMax documents about speech
MiniMax (platform.minimax.io) documents native speech with lip-sync tied to whoever is on screen, and takes reference audio for the voice, capped at 15 seconds across at most three clips. As of 2026-09-12.
| Field | What the vendor documents |
|---|---|
| Audio source | Native speech, generated with the picture |
| Languages | Not documented by the vendor |
| Voice source | Reference audio, 15 s total across 3 clips |
| Per-character binding | By reference audio per call |
| Lip-sync | Tied to the speaker on screen |
Inclusion rule. Fields are the same five for every model here; a field the vendor does not document is recorded as not documented rather than inferred from generated output. Order. Fixed field order, identical on every model page.
1A published figure for how much voice can be supplied
Fifteen seconds in total, across three clips, is the tightest published constraint in this register and the only one stated as a figure rather than a capability. It sets the shape of the work: a voice has to be established from a very short sample, and there is no documented way to store it, so the sample goes in on every call that needs it.
Tying lip-sync to the speaker who is on screen is the other half of the same design. It implies the model is deciding which face moves, which is the part that usually breaks in a two-hander, and it is documented as a behaviour rather than as a control the prompt can override.

2A constraint from elsewhere that lands on dialogue work
A first-frame image and reference images cannot be used in the same call. That restriction is not about audio at all, but it collides with dialogue scenes: continuing from a known frame and carrying a character reference are exactly the two things a conversation between established characters needs at once.
It is recorded here because a register of speech controls that ignored it would describe a workflow the documentation does not permit.
3This entry, one column at a time
Each of these stays inside a single field: what MiniMax puts there, what the wording settles, and what it leaves for a take to answer.
- Audio source — with the picture.
- Voice source — a reference clip on the call.
- Per-character binding — on the call.
- Lip-sync — whoever is on screen.
4Read against another entry
Each of these puts MiniMax beside one other entry on a column where the two land at opposite grades of answer.
- Runway and MiniMax — on voice source.
- MiniMax and Wan 3.0 — on voice source.
- MiniMax and Sora 2 — on lip-sync.
5Where this model sits against the rest
| Model | Audio source |
|---|---|
| D-ID | A text script read aloud, or an audio url supplied |
| Hedra | Supplied; audio is a required input |
| HeyGen | A text to speech endpoint, apart from the picture |
| Kling AI | Generated with the picture, VIDEO 3.0 |
| LTX Studio | Generated jointly with the picture, plus audio-to-video |
| Luma Ray | Not documented; the video reference does not mention sound |
| MiniMax | Native speech, generated with the picture |
| PixVerse | Supplied to a speech endpoint, or read from text there |
| Runway | Not documented where voices are defined |
| SceneMixer | Generated with the picture, in the language set for the project |
| Sora 2 | Generated with the picture; video and audio are both listed as output |
| sync-3 | Supplied, or read from text on the call |
| Synthesia | Synthesised from the script, or an uploaded recording |
| Veo | Generated with the video rather than added afterwards |
| Vidu | Q3 generates speech, effects and music natively, on by default |
| Wan 3.0 | On by default; false returns a file with no audio track |
| Wan2.2-S2V | Supplied; the card describes the model as audio-driven |
Inclusion rule. Models whose vendor documents something about generated speech. A model with nothing published on the point is still listed, with the cell marked hollow. Order. Alphabetical by model name.
| Model | Languages |
|---|---|
| D-ID | A language field beside the voice; no list on the reference |
| Hedra | Full multi-language support claimed, none named |
| HeyGen | Thirty or more, counted on the translation endpoint |
| Kling AI | Five, not named |
| LTX Studio | Not documented |
| Luma Ray | Not documented |
| MiniMax | Not documented by the vendor |
| PixVerse | Multiple, with speech, singing and advertisements named as types |
| Runway | Not documented |
| SceneMixer | 15, named, with Cantonese a 16th for dialogue only |
| Sora 2 | Not documented |
| sync-3 | Ninety-five or more, counted rather than named |
| Synthesia | A catalogue, each voice carrying a language name and code |
| Veo | Not documented |
| Vidu | Not documented |
| Wan 3.0 | Not documented |
| Wan2.2-S2V | Not documented |
Inclusion rule. Models whose vendor documents something about generated speech. A model with nothing published on the point is still listed, with the cell marked hollow. Order. Alphabetical by model name.
| Model | Voice source |
|---|---|
| D-ID | A voice id from one of five named speech providers |
| Hedra | An uploaded track, or a voice id from the voices endpoint |
| HeyGen | A clone, instant from one recording or professional from 20 minutes or more |
| Kling AI | Not documented by the vendor |
| LTX Studio | Not documented |
| Luma Ray | Not documented |
| MiniMax | Reference audio, 15 s total across 3 clips |
| PixVerse | Built-in voices, or a custom voice from a supplied sample |
| Runway | An audio sample of 10 seconds to 5 minutes, or a description in words |
| SceneMixer | Preset voice library, or the user's own recordings |
| Sora 2 | Not documented |
| sync-3 | Not documented on the model page |
| Synthesia | The catalogue, or a cloned voice with a language list of its own |
| Veo | Not documented |
| Vidu | Not documented |
| Wan 3.0 | Reference audio, 15 seconds in total, WAV or MP3 up to 15 MB |
| Wan2.2-S2V | Whatever track is handed over; no catalogue |
Inclusion rule. Models whose vendor documents something about generated speech. A model with nothing published on the point is still listed, with the cell marked hollow. Order. Alphabetical by model name.
6Sources
Taken from the video generation guide at platform.minimax.io on 2026-09-12. Every entry against the same five fields sits on the speech controls page; the rule behind each field is on how read. The entries either side of this one: PixVerse, Runway.
- Audio sourceNative speech with lip-sync tied to the speaker who is on screenspeaker-aware
- Voice sourceReference audio is capped at 15 seconds in total across at most 3 clipsa hard published limit
- Constraint that collides with voice workA first-frame image and reference images cannot be used in the same call