Where a spoken line goes in a prompt
Four entries say where the words go, and no two agree. One wants a labelled block under the prose. One wants the line after a keyword. One wants it quoted. One takes reference voices instead of a written line. As of 2026-09-12.
| Model | What is documented | Detail |
|---|---|---|
| MiniMax | Up to three reference voices | Rather than a written line |
| Sora 2 | A dialogue block under the prose | Speakers labelled, turns alternated |
| Veo | Quote the spoken line | Inside the prompt |
| Wan 3.0 | After the word saying | In the prompt itself |
Inclusion rule. Entries whose documentation shows or states where dialogue is written, whether as a block, a clause, a quotation or a supplied voice. An entry that mentions speech without showing how to ask for it does not earn a row. Order. Alphabetical by model name.
1Dialogue syntax is the part of a shot list that converts worst
A labelled block, a keyword clause and a quotation are not interchangeable. A shot list written against one has to be rewritten for another, and that conversion happens by hand at generation time under deadline, which is exactly where a line gets mangled.
The habit that survives a model change is keeping the line in plain language beside whichever syntax is current. A production storing only rendered prompts has stored the model's format rather than its own script.
2Only one of the four addresses two people talking
Asking for speakers to be labelled consistently, with turns alternated, is the only published guidance here on the commonest shot in drama. The other three describe a single line and leave a conversation to be inferred.
That guidance also converts cleanly into a production's own documents, which is rare. Labelling characters the same way in a shot list is useful whether or not the model reads it, and it means a later change of tool starts from the right shape.
3Supplying a voice instead of writing a line is a different answer
One entry answers this question by taking reference voices rather than text, which puts the words wherever a production already keeps them and the voice on the call. That is not a syntax at all, and it is included here because the practical question is the same: what does the request carry.
It also relocates the risk. A written line is exact and the delivery is a lottery; a reference voice fixes the delivery and leaves the words to whatever the prose implies.
- How a spoken line is written into a promptDialogue belongs in a dialogue block below the prose description, with exchanges limited to a handful of sentences
- Two characters in frameFor multi-character scenes the guide asks for speakers labelled consistently and turns alternated, so each line lands on the right characteraddressed rather than skipped
- How a spoken line is written into a promptThe worked example puts the spoken line in the prompt itself, after the word saying
- Voice sourceReference audio is capped at 15 seconds in total across at most 3 clipsa hard published limit
4Sources
Every line is read from the vendor documentation linked on the speech controls page, checked 2026-09-12. What counts as documented is on how read. Nearby: Auditioning a voice, Outside suppliers named, The column always filled.