Sentioscope

Speech and voice controls, as each vendor documents them

Audio generated with the picture, not after it

Native audio means the sound arrives from the same generation call as the picture, rather than being laid over it afterwards. Four models document it and publish a clip length beside it, and each treats the switch, the billing and the dialogue syntax differently. As of 2026-09-12.

Controls documented, model by modelFilled where the model documents that control, hollow where nothing is published about it.Controls documented, model by modelLength per callHow audio is pricedHow audio ispricedHow dialogue is writtenHow dialogue iswrittenMiniMax H3MiniMax H3 — Length per call: 4 to 15 sMiniMax H3 — How audio is priced: No separate audio line item publishedMiniMax H3 — How dialogue is written: Up to three reference voicesSora 2Sora 2 — Length per call: 16 or 20 s, extendable to 120Sora 2 — How audio is priced: Priced per second of video generatedSora 2 — How dialogue is written: A dialogue block under the proseVeo 3.1Veo 3.1 — Length per call: 4, 6 or 8 sVeo 3.1 — How audio is priced: Included in the per-second priceVeo 3.1 — How dialogue is written: Quote the spoken lineWan 3.0Wan 3.0 — Length per call: 2 to 30 sWan 3.0 — How audio is priced: On by default; switching it off does not reduce the priceWan 3.0 — How dialogue is written: Write the line after the word saying
Fig. 1 Filled where the model documents that control, hollow where nothing is published about it.
Native audio, as published. Recorded 2026-09-12.
ModelLength per callHow audio is pricedHow dialogue is written
MiniMax H34 to 15 sNo separate audio line item publishedUp to three reference voices
Sora 216 or 20 s, extendable to 120Priced per second of video generatedA dialogue block under the prose
Veo 3.14, 6 or 8 sIncluded in the per-second priceQuote the spoken line
Wan 3.02 to 30 sOn by default; switching it off does not reduce the priceWrite the line after the word saying

Inclusion rule. Models whose vendor documents audio produced in the same call as the picture and publishes a clip length to write against. A model where sound is a separate product, or where no length is published, is not given a row. Order. Alphabetical by model name.

1Off is not cheaper

One vendor states plainly that turning audio off does not change the rate. That makes silence a creative decision rather than a budget one, and it means a production generating silent picture is paying for a capability it discards.

Where audio is included in the price, the useful question is whether it is good enough to keep. If it is not, the cost of replacing it lands entirely outside the rate card.

2Silence has to be asked for

With audio on by default, a shot that should be quiet will not be unless the prompt says so. One vendor documents the exact phrasing for this, which is unusual and worth copying: an explicit negative rather than an absence.

Productions that forget this get music under scenes that were meant to play dry, and the fix is a regeneration rather than a mute in the edit, because the music is baked into the returned file.

3Dialogue syntax is per model and does not transfer

One vendor wants the line in quotation marks, another takes reference voices, another expects explicit negatives around it. A shot list carrying dialogue therefore has to be written against the model that will render it.

That is the part of a shot list that converts least cleanly, and the part most worth keeping in plain language alongside whatever syntax is current.

4What the audio is good enough for

Generated speech arriving with the picture is aligned by construction, which is the hard part. What it usually is not is a finished mix: levels move between calls, room tone changes, and music comes and goes with the prompt.

Treating it as a guide track that happens to be in sync, rather than as delivery audio, is the safer assumption and the one that keeps a post budget honest.

5Regenerating a shot regenerates the voice

Because the audio comes from the same call, redoing a shot for a visual reason produces a new performance as well. Where a line was already approved, that is a cost nobody planned, and it is the reason productions separate picture fixes from audio fixes wherever the tool allows it.

6Sources read for this entry

This page defines the term and logs what each vendor documents about it, each figure read on 2026-09-12. Related: Reference audio, Dialogue timing.