Sentioscope

Speech and voice controls, as each vendor documents them

Which models return sound unless told otherwise

Wan 3.0 defaults the audio parameter to true and states that enabling or disabling audio does not affect pricing. Vidu's Q3 generates speech, effects and music with audio on by default and documents a speech-only mode. Sora 2 lists audio as an output beside video. As of 2026-09-12.

What arrives with the picture when nothing is specifiedOne model generates speech, sound effects and music natively with audio on by default in its API and documents a speech-only mode. Another generates audio with the video rather than adding it afterwards. A third documents joint generation plus a route that runs from audio to video.A call with nothing specifiedAudio on by defaultSpeech, effects andmusicA speech-only mode isdocumented separately.Generated jointlyAudio arrives withthe pictureNot added as a later pass.Audio to videoThe pipeline runsbackwardsThe sound leads and thepicture follows.The default is what most output ends up sounding like
Fig. 1 A default decides what most output sounds like, because most calls do not override it. Speech-only is the mode a drama pipeline actually wants.
What each model returns when the caller does not specify anything about audio. Recorded 2026-09-12.
ModelWhat is documentedDetail
Wan 3.0On by defaultFalse returns no audio track, and the price does not move
Vidu Q3On by defaultSpeech, effects and music; speech-only mode available
Sora 2With the pictureVideo and audio are both listed as output
VeoWith the videoGenerated rather than added afterwards
LTX StudioJoint generationPlus an audio-to-video direction

Inclusion rule. Entries whose documentation states a default or describes joint generation. An entry describing audio without saying whether it is default is recorded as not stated. Order. Entries stating an explicit default first.

1A default decides what most output sounds like

Every call made without reading the parameter list returns whatever the default produces. For one model here that is a finished soundscape rather than a silent picture, which means its baseline output is not comparable with a model that returns silence unless asked.

Anyone benchmarking two models without checking this is comparing different products. It is a one-line fact and it changes what a side-by-side test is actually measuring.

2Speech-only is the mode a drama pipeline wants

Separating the line from the effects and the music means dialogue can be kept while the bed is replaced by a composer or a sound editor. One model documents that separation; nothing else in this register does.

For a serial production that is the difference between generated audio being an asset and being something to work around, and it is a documentation fact rather than a claim about how the audio sounds.

3Audio-to-video runs the pipeline backwards

One entry documents driving generation from an existing track rather than producing one. For music and advertising that inverts the usual order: the audio is fixed and the picture is fitted to it.

What it means for drama is not addressed. A recorded dialogue track is an existing track, and whether the route is intended to accept one is not stated anywhere public, so the register records the capability and not a workflow.

  • Audio source
    Q3 outputs speech, sound effects and music natively, on by default in the API, with a speech-only modethree tracks in one passVidu, API reference / recorded 2026-09-12
  • Audio source
    Audio is generated with the video rather than added afterwardsgenerated with the pictureGoogle, Veo documentation / recorded 2026-09-12
  • Audio source
    Joint audio and video generation, plus audio-to-video, on LTX-2.5two directionsLTX Studio / recorded 2026-09-12
  • Audio source
    The audio parameter defaults to true so the returned video carries an audio track, and false returns one withouton unless it is switched offAlibaba Cloud, Wan 3.0 API reference / recorded 2026-09-22
  • What choosing silence costs
    Enabling or disabling audio does not affect pricingAlibaba Cloud, Wan 3.0 API reference / recorded 2026-09-22
  • Audio source
    Described as a media generation model generating videos with synced audio, with video and audio both listed as outputgenerated with the pictureOpenAI, Sora 2 model page / recorded 2026-09-22

4Sources

Every line is read from the vendor documentation linked on the speech controls page, checked 2026-09-12. What counts as documented is on how read. Nearby: Which language, Two in frame, Who docs are for.