Sentioscope

Speech and voice controls, as each vendor documents them

Entries that rebuild the voice on every generation

Five entries put the voice on the request. Two pass audio, which can drift; three pass an identifier or a label, which cannot. In all five the pairing between a character and a voice lives in a production's own records. As of 2026-09-22.

A clip and an identifier are not the same per-callAn identifier is short, stable and impossible to drift. A reference clip has to be the same file every time and nothing checks that it is, so one arrangement produces two very different risk profiles.What actually travels on the requestAn id or a labelCannot driftShort, stable, and easy to keep in a shotlistA reference clipCan drift silentlyNothing checks it is the same file, andnothing records which one spokeIn both cases the pairing lives outside the platform
Fig. 1 Fifteen seconds in total means a scene with three speaking characters divides the allowance three ways.
Entries that re-establish a voice per call, and what each one actually passes. Recorded 2026-09-22.
ModelWhat travels on the requestAnd where it comes from
D-IDOn the requestA voice id, or a recording by url
MiniMaxOn the callA reference clip on the call
PixVerseOn the requestA sample, or a built-in voice
Sora 2On the promptNothing documented as an input
Wan 3.0On the callReference audio, 15 seconds in total

Inclusion rule. Entries whose documentation describes the voice being specified on each generation, whether as audio, an identifier or a label in the prompt. Entries with a stored voice are on a separate route page. Order. Alphabetical by model name.

1A clip and an identifier are not the same kind of per-call

An id is short, stable and impossible to drift, so a production that writes it down gets consistency for free. A reference clip has to be the same file every time, and nothing checks that it is, so the same arrangement produces two very different risk profiles.

Both still leave the bookkeeping outside the platform. Neither a clip nor an id is attached to a character by anything the vendor holds, which means a shot list is the authoritative record of a cast.

2Two entries publish an identical fifteen-second ceiling

Fifteen seconds in total is the figure two of these publish for reference audio, and the arrangements around it differ: one divides the total across at most three clips, the other lets a single clip fill it. That decides whether a cast is supplied as one clean take or several short ones.

Fifteen seconds also means a scene with three speaking characters divides the allowance three ways, so voices go in one per shot. The call cannot hold a conversation, which is a constraint on coverage as much as on casting.

3A label is the weakest member of this group

One entry passes neither audio nor an id but a speaker name inside the prompt. Within one generation that routes each line to the right face, which is real work most vendors leave to chance. Across generations it does nothing at all.

Grouping it here rather than under nothing published is deliberate: something is being specified per call, and a reader deciding how to plan a two-hander needs to know it exists. What it cannot do is substitute for a cast.

4The entries on this route, one page each

Each of these links to that entry on the field this route groups by. The wording behind every cell, and the date it was read, sits on the page it links to.

  • Voice source
    Reference audio is capped at 15 seconds in total across at most 3 clipsa hard published limitMiniMax, video generation guide / recorded 2026-09-12
  • Voice source
    Reference audio is WAV or MP3, one to fifteen seconds a clip, fifteen seconds in total and no more than 15 MBa published capAlibaba Cloud, Wan 3.0 API reference / recorded 2026-09-22
  • Per-character binding
    The chosen speaker id is passed as the lip-sync speaker on the generation requesta speaker id per callPixVerse, speech and lip sync guide / recorded 2026-09-22
  • Per-character binding
    The voice is a voice id selected from the list of available voices, with an optional language field beside ita voice id on every requestD-ID, create a talk reference / recorded 2026-09-22
  • Two characters in frame
    For multi-character scenes the guide asks for speakers labelled consistently and turns alternated, so each line lands on the right characteraddressed rather than skippedOpenAI, Sora 2 prompting guide / recorded 2026-09-22
  • Constraint that collides with voice work
    A first-frame image and reference images cannot be used in the same callMiniMax, video generation guide / recorded 2026-09-12

5Sources

Membership of this route is decided by the wording in the field it groups by, read from the documentation each vendor publishes. The column itself is on the field note, all five columns are on the speech table, and what counts as documented is on how read. Other routes: Languages named, Languages counted.