Audio generated with the picture, not after it
Native audio means the sound arrives from the same generation call as the picture, rather than being laid over it afterwards. Four models document it and publish a clip length beside it, and each treats the switch, the billing and the dialogue syntax differently. As of 2026-09-12.
| Model | Length per call | How audio is priced | How dialogue is written |
|---|---|---|---|
| MiniMax H3 | 4 to 15 s | No separate audio line item published | Up to three reference voices |
| Sora 2 | 16 or 20 s, extendable to 120 | Priced per second of video generated | A dialogue block under the prose |
| Veo 3.1 | 4, 6 or 8 s | Included in the per-second price | Quote the spoken line |
| Wan 3.0 | 2 to 30 s | On by default; switching it off does not reduce the price | Write the line after the word saying |
Inclusion rule. Models whose vendor documents audio produced in the same call as the picture and publishes a clip length to write against. A model where sound is a separate product, or where no length is published, is not given a row. Order. Alphabetical by model name.
1Off is not cheaper
One vendor states plainly that turning audio off does not change the rate. That makes silence a creative decision rather than a budget one, and it means a production generating silent picture is paying for a capability it discards.
Where audio is included in the price, the useful question is whether it is good enough to keep. If it is not, the cost of replacing it lands entirely outside the rate card.
2Silence has to be asked for
With audio on by default, a shot that should be quiet will not be unless the prompt says so. One vendor documents the exact phrasing for this, which is unusual and worth copying: an explicit negative rather than an absence.
Productions that forget this get music under scenes that were meant to play dry, and the fix is a regeneration rather than a mute in the edit, because the music is baked into the returned file.
3Dialogue syntax is per model and does not transfer
One vendor wants the line in quotation marks, another takes reference voices, another expects explicit negatives around it. A shot list carrying dialogue therefore has to be written against the model that will render it.
That is the part of a shot list that converts least cleanly, and the part most worth keeping in plain language alongside whatever syntax is current.
4What the audio is good enough for
Generated speech arriving with the picture is aligned by construction, which is the hard part. What it usually is not is a finished mix: levels move between calls, room tone changes, and music comes and goes with the prompt.
Treating it as a guide track that happens to be in sync, rather than as delivery audio, is the safer assumption and the one that keeps a post budget honest.
5Regenerating a shot regenerates the voice
Because the audio comes from the same call, redoing a shot for a visual reason produces a new performance as well. Where a line was already approved, that is a cost nobody planned, and it is the reason productions separate picture fixes from audio fixes wherever the tool allows it.
6Sources read for this entry
This page defines the term and logs what each vendor documents about it, each figure read on 2026-09-12. Related: Reference audio, Dialogue timing.