Sentioscope

Speech and voice controls, as each vendor documents them

Wan 3.0: on unless switched off, and priced the same

The audio parameter defaults to true, so a returned video carries a track unless the call says false. The reference adds the part that changes a budget conversation: enabling or disabling audio does not affect pricing. As of 2026-09-22.

A switch with no price attached to either positionThe audio parameter defaults to true, and the reference states that enabling or disabling audio does not affect pricing. Generating silent picture therefore means paying for a capability and discarding it.The audio parameter on a Wan 3.0 callLeft at its default, trueA file with a trackMusic and effects arrive unless theprompt asks for silenceSet to falseA file with no trackSame rate, so the only reason is notwanting the soundPricing is stated to be identical either way
Fig. 1 Turned around, the same sentence says the audio is free to try on every shot, so whether it is worth keeping becomes the whole question.
Wan 3.0 on audio source, statement by statement. Read from the vendor's video generation API reference on 2026-09-22.
What the documentation settlesWhat it leaves to a take
The audio parameter defaults to true; false returns a file with no audio trackWhat the default track contains when the prompt asks for nothing
Enabling or disabling audio does not affect pricingWhether the generated audio is good enough to keep at that rate
The worked example puts the spoken line in the prompt, after the word sayingHow several speakers in one shot are written in that syntax

Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.

1Silence costs the same, which makes it a creative decision

Most audio features are sold as an upgrade, so a team that does not want the sound expects to save by switching it off. Here the vendor says plainly that the rate does not move. Generating silent picture therefore means paying for a capability and discarding it, and the only reason left to do so is that the sound is not wanted.

Turned the other way, the same sentence says the audio is free to try on every shot. Whether it is good enough to keep is then the whole question, and the cost of replacing it lands entirely outside the rate card.

2The line goes in the prompt, beside the camera direction

Putting the spoken words after the word saying, inside the same prose that describes the shot, is a syntax rather than a parameter. It means the dialogue and the staging are edited in the same string, and a shot list has to carry them together or convert them at generation time.

That is the part of a shot list that travels least well between models. One vendor wants a labelled block, another takes reference voices, this one wants a clause. Keeping the line in plain language alongside whichever syntax is current is the habit that survives a model change.

3The others that make sound while they make the picture

Seven more entries generate audio in the same pass. This is the only one that states what choosing silence costs.

  • Kling AI — with the picture, on video 3.0.
  • LTX Studio — with the picture, and audio to video.
  • MiniMax — with the picture.
  • SceneMixer — with the picture, in the language set for the project.
  • Sora 2 — with the picture.
  • Veo — with the picture.
  • Vidu — with the picture, speech available on its own.

4Sources

Read from the video generation API reference at alibabacloud.com on 2026-09-22. The same column across every entry is on audio source; everything this vendor publishes about speech is on Wan 3.0. What counts as documented is on how read.