Sentioscope

Speech and voice controls, as each vendor documents them

Where a spoken line goes in a prompt

Four entries say where the words go, and no two agree. One wants a labelled block under the prose. One wants the line after a keyword. One wants it quoted. One takes reference voices instead of a written line. As of 2026-09-12.

Four places a spoken line is supposed to goA labelled block under the prose, a clause after one keyword, a quotation inside the prompt, or reference voices instead of written words. No two of the four are interchangeable.A labelled blockUnder the proseSpeakers namedTurns can alternateAfter a keywordInside the promptOne clauseStaging and wordstogetherA quotationInside the promptMarked wordsDeliveryunspecifiedReference voicesOn the callUp to threeNo written line atallOne generation carrying speechConversion between these is where a line gets mangled
Fig. 1 A production storing only rendered prompts has stored the model's format rather than its own script.
How each vendor documents writing a spoken line into a request. Recorded 2026-09-12.
ModelWhat is documentedDetail
MiniMaxUp to three reference voicesRather than a written line
Sora 2A dialogue block under the proseSpeakers labelled, turns alternated
VeoQuote the spoken lineInside the prompt
Wan 3.0After the word sayingIn the prompt itself

Inclusion rule. Entries whose documentation shows or states where dialogue is written, whether as a block, a clause, a quotation or a supplied voice. An entry that mentions speech without showing how to ask for it does not earn a row. Order. Alphabetical by model name.

1Dialogue syntax is the part of a shot list that converts worst

A labelled block, a keyword clause and a quotation are not interchangeable. A shot list written against one has to be rewritten for another, and that conversion happens by hand at generation time under deadline, which is exactly where a line gets mangled.

The habit that survives a model change is keeping the line in plain language beside whichever syntax is current. A production storing only rendered prompts has stored the model's format rather than its own script.

2Only one of the four addresses two people talking

Asking for speakers to be labelled consistently, with turns alternated, is the only published guidance here on the commonest shot in drama. The other three describe a single line and leave a conversation to be inferred.

That guidance also converts cleanly into a production's own documents, which is rare. Labelling characters the same way in a shot list is useful whether or not the model reads it, and it means a later change of tool starts from the right shape.

3Supplying a voice instead of writing a line is a different answer

One entry answers this question by taking reference voices rather than text, which puts the words wherever a production already keeps them and the voice on the call. That is not a syntax at all, and it is included here because the practical question is the same: what does the request carry.

It also relocates the risk. A written line is exact and the delivery is a lottery; a reference voice fixes the delivery and leaves the words to whatever the prose implies.

  • How a spoken line is written into a prompt
    Dialogue belongs in a dialogue block below the prose description, with exchanges limited to a handful of sentencesOpenAI, Sora 2 prompting guide / recorded 2026-09-22
  • Two characters in frame
    For multi-character scenes the guide asks for speakers labelled consistently and turns alternated, so each line lands on the right characteraddressed rather than skippedOpenAI, Sora 2 prompting guide / recorded 2026-09-22
  • How a spoken line is written into a prompt
    The worked example puts the spoken line in the prompt itself, after the word sayingAlibaba Cloud, Wan 3.0 API reference / recorded 2026-09-22
  • Voice source
    Reference audio is capped at 15 seconds in total across at most 3 clipsa hard published limitMiniMax, video generation guide / recorded 2026-09-12

4Sources

Every line is read from the vendor documentation linked on the speech controls page, checked 2026-09-12. What counts as documented is on how read. Nearby: Auditioning a voice, Outside suppliers named, The column always filled.