Sora 2 and Wan 3.0: two ways to write a spoken line
Two native-audio entries that tell a writer where the words go. One asks for a dialogue block underneath the prose description with speakers labelled. The other puts the line in the prompt itself, after the word saying. As of 2026-09-22.
| Field | Sora 2 | Wan 3.0 | Where they part |
|---|---|---|---|
| Audio source | With the picture | With the picture, unless switched off | Different answers |
| Voice source | Nothing documented as an input | Reference audio, 15 seconds in total | Different answers |
| Per-character binding | On the prompt | On the call | Different answers |
| Languages | Not documented by the vendor | Not documented by the vendor | Same answer |
| Lip-sync | Named only where it fails | Not documented by the vendor | Only Sora 2 answers |
Inclusion rule. Two entries are given a page together when at least one column puts them at opposite grades of answer. Pairs that agree on every column, or that are both blank throughout, do not get a page. Order. Fixed field order, identical on every side-by-side page.
1Dialogue syntax is the part of a shot list that travels worst
A labelled block and an inline clause are not interchangeable, and a shot list written for one has to be converted for the other. That conversion is exactly where a line gets mangled, because it is done by hand at generation time under deadline.
The habit that survives a model change is keeping the line in plain language alongside whichever syntax is current. A production that stores only the rendered prompt has stored the model's format rather than its own script.
2Both publish numbers, and they bound different things
One publishes clip lengths of sixteen or twenty seconds extendable to a hundred and twenty, plus how much speech fits four and eight seconds. The other publishes a default, a pricing statement, reference audio limits and a byte ceiling.
Between them a writer can size a scene on one and cost a voice on the other, and neither page carries both halves. Reading them together is how the arithmetic gets finished.
3One addresses two speakers, the other addresses silence
One asks for turns alternated so each line lands on the right face, which is the commonest shot in drama. The other documents the phrasing for asking a shot to be quiet, since audio is on by default and silence costs the same.
Both are unusual sentences, and neither vendor publishes the other one. A production needing both has to take one from each page and treat the combination as its own convention.
4Each of them on its own
The column this pair was chosen for is audio source, and each entry has a page of its own on it. The full row for either, all five columns with the wording behind each cell, is on its model note.
- Sora 2 on audio source — with the picture.
- Wan 3.0 on audio source — with the picture, unless switched off.
- Sora 2, all five fields — read from the model page and prompting guide.
- Wan 3.0, all five fields — read from the video generation API reference.
- How a spoken line is written into a promptDialogue belongs in a dialogue block below the prose description, with exchanges limited to a handful of sentences
- Shot length a line has to fitGenerations run 16 or 20 seconds, and an extension adds up to 20 seconds at a time to a total of 120
- Two characters in frameFor multi-character scenes the guide asks for speakers labelled consistently and turns alternated, so each line lands on the right characteraddressed rather than skipped
- How a spoken line is written into a promptThe worked example puts the spoken line in the prompt itself, after the word saying
- What choosing silence costsEnabling or disabling audio does not affect pricing
5Sources
Every cell above is read from the documentation each vendor publishes, on the dates carried by the two model notes. The whole register on one column is on the field note; entries grouped by the shape of their answer are on the routes. Other pairs: Runway and Luma Ray, LTX Studio and Vidu.