Sentioscope

Speech and voice controls, as each vendor documents them

Sora 2: a label routes a line, and then it is gone

Speaker labels inside the prompt send each line to the right face. A label is an instruction within one call rather than an identity that survives the next, so nothing here makes a character sound the same in the following episode. As of 2026-09-22.

Inside one call a label works, and then it stopsA speaker label tells the model which of two people is speaking, which most vendors leave to chance. Across generations it does nothing, because no mechanism connects a label to a voice.Within one generationAcross generationsWhat a label doesRoutes each line to the right faceNothing publishedWhat the guide asksConsistent labels, turns alternatedNothing about reuseWhat a voice isUnspecified, but internally consistentUnspecified, and unconnectedWorth doing anywayYes, the model behaves predictablyYes, the documents stay usableLabels solve the shot and not the series
Fig. 1 A two-hander can play correctly in one shot and come back with different voices in the reverse angle.
Sora 2 on per-character binding, statement by statement. Read from the vendor's model page and prompting guide on 2026-09-22.
What the documentation settlesWhat it leaves to a take
Speakers should be labelled consistently, with turns alternatedWhether the same label produces the same voice twice
Dialogue belongs in a dialogue block below the prose descriptionWhether a label can be tied to anything outside that block
Nothing is published about where a voice comes fromWhether a voice can be chosen, supplied or reused at all

Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.

1Labels solve the shot and not the series

Inside one generation a label does real work: it tells the model which of two people is speaking, which is the distinction most vendors leave to chance. The guidance is careful about it, asking for consistency and alternating turns.

Across generations it does nothing, because there is no stated mechanism connecting a label to a voice. A two-hander can play correctly in one shot and come back with different voices in the reverse angle.

2Consistent labelling is still worth doing

Even without persistence, labelling the same character the same way keeps a shot list readable and keeps the model's behaviour inside one call predictable. It also means that if a voice mechanism ever arrives, the production's documents are already in the shape it would want.

What it cannot do is substitute for a cast. A serial needs voices that outlast a call, and this entry publishes nothing that does, which is the finding rather than the workaround.

3The others that re-establish the voice on each call

Four more entries put the voice on the request. All four pass something that identifies a voice; this one passes a name that identifies a speaker and leaves the voice unaddressed.

  • Two characters in frame
    For multi-character scenes the guide asks for speakers labelled consistently and turns alternated, so each line lands on the right characteraddressed rather than skippedOpenAI, Sora 2 prompting guide / recorded 2026-09-22
  • How a spoken line is written into a prompt
    Dialogue belongs in a dialogue block below the prose description, with exchanges limited to a handful of sentencesOpenAI, Sora 2 prompting guide / recorded 2026-09-22
  • Languages
    Not documented by the vendor (as of 2026-09-22)neither a list nor a countOpenAI, video generation guide / recorded 2026-09-22

4Sources

Read from the model page and prompting guide at developers.openai.com on 2026-09-22. The same column across every entry is on per-character binding; everything this vendor publishes about speech is on Sora 2. What counts as documented is on how read.