Sentioscope

Speech and voice controls, as each vendor documents them

Whether generated audio is charged for separately

Nothing in this register charges for audio on its own. Where cost is mentioned at all the sound sits inside a per-second rate for the video, or has no line of its own, and one vendor states that disabling it does not reduce the price. As of 2026-09-12.

No price on the sound, and no reason to describe itIf audio has no price of its own, no vendor has an incentive to describe its quality and a production has no way to weigh it against an alternative. The sound arrives, and whether it is usable is discovered.Bundled inAudio inside aper-second videorate, or with no lineof its own.SoUnpricedNo incentive todescribe quality, andno basis forcomparison.SoDiscoveredA production findsout by generating,and budgets a sessionin case.
Fig. 1 Budget for replacing it and be pleased if the budget goes unspent; that keeps a post-production estimate truthful.
Controls documented, model by modelFilled where the model documents that control, hollow where nothing is published about it.Controls documented, model by modelWhat is documentedMiniMaxMiniMax — What is documented: No separate audio line publishedSora 2Sora 2 — What is documented: Priced per second of video generatedVeoVeo — What is documented: Included in the per-second priceWan 3.0Wan 3.0 — What is documented: Enabling or disabling audio does not affect pricing
Fig. 2 Filled where the model documents that control, hollow where nothing is published about it.
What each vendor publishes about paying for the audio it generates. Recorded 2026-09-12.
ModelWhat is documentedDetail
MiniMaxNo separate audio line publishedNative speech with the picture
Sora 2Priced per second of video generatedAudio is part of that generation
VeoIncluded in the per-second priceGenerated with the video
Wan 3.0Enabling or disabling audio does not affect pricingStated in the reference

Inclusion rule. Entries whose documentation says anything about how generated audio is priced. Entries publishing a rate for video alone, with no statement about the audio inside it, do not earn a row. Order. Alphabetical by model name.

1Bundling sound into a video rate has a quiet consequence

If the audio has no price of its own, then no vendor has an incentive to describe its quality, and a production has no way to weigh it against an alternative. The sound arrives, it is free at the margin, and whether it is usable is discovered rather than priced.

The honest way to plan around that is to budget for replacing it and be pleased if the budget goes unspent. Treating generated dialogue as a guide track keeps a post-production estimate truthful.

2The real cost of audio is outside every rate card here

Replacing a generated performance means a session, a performer, a mix and possibly a regenerated shot. None of that appears on a pricing page, and it is the largest number in the conversation for any series that decides the generated read is not good enough.

That is why a statement about what silence costs is worth recording even though it looks trivial. It tells a reader the vendor has thought about the switch, and it removes a false saving from a plan.

3Most of the register is silent, because the documents are technical

Thirteen entries say nothing about cost at all. The pages read for this register are references, cards and guides rather than price lists, so commercial statements appear incidentally if they appear.

Recording the four that say something is still the right thing to do. Incidental sentences about pricing are usually more candid than a pricing page, because nobody wrote them to be persuasive.

  • What choosing silence costs
    Enabling or disabling audio does not affect pricingAlibaba Cloud, Wan 3.0 API reference / recorded 2026-09-22
  • Shot length a line has to fit
    Generations run 16 or 20 seconds, and an extension adds up to 20 seconds at a time to a total of 120OpenAI, video generation guide / recorded 2026-09-22
  • Audio source
    Audio is generated with the video rather than added afterwardsgenerated with the pictureGoogle, Veo documentation / recorded 2026-09-12
  • Audio source
    Native speech with lip-sync tied to the speaker who is on screenspeaker-awareMiniMax, video generation guide / recorded 2026-09-12
  • Audio source
    The audio parameter defaults to true so the returned video carries an audio track, and false returns one withouton unless it is switched offAlibaba Cloud, Wan 3.0 API reference / recorded 2026-09-22

4Sources

Every line is read from the vendor documentation linked on the speech controls page, checked 2026-09-12. What counts as documented is on how read. Nearby: Keeping a voice, Lip-sync, Voice sources.