Entries that rebuild the voice on every generation
Five entries put the voice on the request. Two pass audio, which can drift; three pass an identifier or a label, which cannot. In all five the pairing between a character and a voice lives in a production's own records. As of 2026-09-22.
| Model | What travels on the request | And where it comes from |
|---|---|---|
| D-ID | On the request | A voice id, or a recording by url |
| MiniMax | On the call | A reference clip on the call |
| PixVerse | On the request | A sample, or a built-in voice |
| Sora 2 | On the prompt | Nothing documented as an input |
| Wan 3.0 | On the call | Reference audio, 15 seconds in total |
Inclusion rule. Entries whose documentation describes the voice being specified on each generation, whether as audio, an identifier or a label in the prompt. Entries with a stored voice are on a separate route page. Order. Alphabetical by model name.
1A clip and an identifier are not the same kind of per-call
An id is short, stable and impossible to drift, so a production that writes it down gets consistency for free. A reference clip has to be the same file every time, and nothing checks that it is, so the same arrangement produces two very different risk profiles.
Both still leave the bookkeeping outside the platform. Neither a clip nor an id is attached to a character by anything the vendor holds, which means a shot list is the authoritative record of a cast.
2Two entries publish an identical fifteen-second ceiling
Fifteen seconds in total is the figure two of these publish for reference audio, and the arrangements around it differ: one divides the total across at most three clips, the other lets a single clip fill it. That decides whether a cast is supplied as one clean take or several short ones.
Fifteen seconds also means a scene with three speaking characters divides the allowance three ways, so voices go in one per shot. The call cannot hold a conversation, which is a constraint on coverage as much as on casting.
3A label is the weakest member of this group
One entry passes neither audio nor an id but a speaker name inside the prompt. Within one generation that routes each line to the right face, which is real work most vendors leave to chance. Across generations it does nothing at all.
Grouping it here rather than under nothing published is deliberate: something is being specified per call, and a reader deciding how to plan a two-hander needs to know it exists. What it cannot do is substitute for a cast.
4The entries on this route, one page each
Each of these links to that entry on the field this route groups by. The wording behind every cell, and the date it was read, sits on the page it links to.
- D-ID — on the request.
- MiniMax — on the call.
- PixVerse — on the request.
- Sora 2 — on the prompt.
- Wan 3.0 — on the call.
- Voice sourceReference audio is capped at 15 seconds in total across at most 3 clipsa hard published limit
- Voice sourceReference audio is WAV or MP3, one to fifteen seconds a clip, fifteen seconds in total and no more than 15 MBa published cap
- Per-character bindingThe chosen speaker id is passed as the lip-sync speaker on the generation requesta speaker id per call
- Per-character bindingThe voice is a voice id selected from the list of available voices, with an optional language field beside ita voice id on every request
- Two characters in frameFor multi-character scenes the guide asks for speakers labelled consistently and turns alternated, so each line lands on the right characteraddressed rather than skipped
- Constraint that collides with voice workA first-frame image and reference images cannot be used in the same call
5Sources
Membership of this route is decided by the wording in the field it groups by, read from the documentation each vendor publishes. The column itself is on the field note, all five columns are on the speech table, and what counts as documented is on how read. Other routes: Languages named, Languages counted.