Sentioscope

Speech and voice controls, as each vendor documents them

What it costs to ask for a shot with no sound

Audio is usually sold as an addition, so a team that does not want sound expects to save by switching it off. One vendor says the rate does not move. Two others fold audio into a per-second price, which comes to the same thing. As of 2026-09-12.

A switch that saves nothing, read two waysIf the rate does not move, generating silent picture means paying for a capability and discarding it. Turned around, the same sentence says the audio is free to attempt on every shot.Audio off, on a rate that does not changeRead as a costPaying for nothingA silent deliverable at the price of asounded oneRead as an offerFree to tryEvery shot can be heard, and kept only ifit is usableThirteen entries say nothing about this at all
Fig. 1 The real cost of speech lands outside the rate card, in the session that replaces a read nobody kept.
Controls documented, model by modelFilled where the model documents that control, hollow where nothing is published about it.Controls documented, model by modelWhat is documentedMiniMaxMiniMax — What is documented: No separate audio line publishedSora 2Sora 2 — What is documented: Priced per second of video generatedVeoVeo — What is documented: Included in the per-second priceWan 3.0Wan 3.0 — What is documented: Enabling or disabling audio does not affect pricing
Fig. 2 Filled where the model documents that control, hollow where nothing is published about it.
What each vendor publishes about paying for audio it generates. Recorded 2026-09-12.
ModelWhat is documentedDetail
MiniMaxNo separate audio line publishedFour to fifteen seconds per generation
Sora 2Priced per second of video generatedAudio is part of the generation
VeoIncluded in the per-second priceFour, six or eight second shots
Wan 3.0Enabling or disabling audio does not affect pricingStated in the reference

Inclusion rule. Entries whose documentation says something about how generated audio is charged for, including saying that switching it off changes nothing. Entries silent on cost do not earn a row. Order. Alphabetical by model name.

1Free to try is not the same as free to keep

If the rate is identical either way, then generated audio costs nothing to attempt on every shot, and the only question left is whether it is good enough to keep. That is a better position than paying extra for it, and it moves the whole cost of speech outside the rate card.

Because replacing it is where the money goes. A production that discards generated dialogue is paying for a capability it throws away and then paying again for a session, and neither half of that appears on a pricing page.

2A default that costs nothing extra is a default nobody turns off

Where audio is on by default and switching it off saves nothing, most output from that model will carry sound. Teams benchmarking it against a silent generator are comparing two different deliverables, and the one that arrives with music usually feels further along than it is.

It also means silence has to be asked for explicitly. A scene meant to play dry will not unless the prompt says so, and the fix afterwards is a regeneration rather than a mute, because the bed is inside the file.

3Thirteen entries say nothing at all about this

Cost is the field a production asks about first and the one this register can say least about, because most documentation read here is technical rather than commercial. Where a sentence exists it is usually incidental, as in a note that a parameter does not affect pricing.

Recording the four that say anything is still worth doing. A vendor willing to state what a switch costs has told a reader something about how it thinks about the capability.

  • What choosing silence costs
    Enabling or disabling audio does not affect pricingAlibaba Cloud, Wan 3.0 API reference / recorded 2026-09-22
  • Audio source
    The audio parameter defaults to true so the returned video carries an audio track, and false returns one withouton unless it is switched offAlibaba Cloud, Wan 3.0 API reference / recorded 2026-09-22
  • Shot length a line has to fit
    Generations run 16 or 20 seconds, and an extension adds up to 20 seconds at a time to a total of 120OpenAI, video generation guide / recorded 2026-09-22
  • Audio source
    Audio is generated with the video rather than added afterwardsgenerated with the pictureGoogle, Veo documentation / recorded 2026-09-12
  • Audio source
    Native speech with lip-sync tied to the speaker who is on screenspeaker-awareMiniMax, video generation guide / recorded 2026-09-12

4Sources

Every line is read from the vendor documentation linked on the speech controls page, checked 2026-09-12. What counts as documented is on how read. Nearby: Where a line goes, Auditioning a voice, Outside suppliers named.