Which models return sound unless told otherwise
Wan 3.0 defaults the audio parameter to true and states that enabling or disabling audio does not affect pricing. Vidu's Q3 generates speech, effects and music with audio on by default and documents a speech-only mode. Sora 2 lists audio as an output beside video. As of 2026-09-12.
| Model | What is documented | Detail |
|---|---|---|
| Wan 3.0 | On by default | False returns no audio track, and the price does not move |
| Vidu Q3 | On by default | Speech, effects and music; speech-only mode available |
| Sora 2 | With the picture | Video and audio are both listed as output |
| Veo | With the video | Generated rather than added afterwards |
| LTX Studio | Joint generation | Plus an audio-to-video direction |
Inclusion rule. Entries whose documentation states a default or describes joint generation. An entry describing audio without saying whether it is default is recorded as not stated. Order. Entries stating an explicit default first.
1A default decides what most output sounds like
Every call made without reading the parameter list returns whatever the default produces. For one model here that is a finished soundscape rather than a silent picture, which means its baseline output is not comparable with a model that returns silence unless asked.
Anyone benchmarking two models without checking this is comparing different products. It is a one-line fact and it changes what a side-by-side test is actually measuring.
2Speech-only is the mode a drama pipeline wants
Separating the line from the effects and the music means dialogue can be kept while the bed is replaced by a composer or a sound editor. One model documents that separation; nothing else in this register does.
For a serial production that is the difference between generated audio being an asset and being something to work around, and it is a documentation fact rather than a claim about how the audio sounds.
3Audio-to-video runs the pipeline backwards
One entry documents driving generation from an existing track rather than producing one. For music and advertising that inverts the usual order: the audio is fixed and the picture is fitted to it.
What it means for drama is not addressed. A recorded dialogue track is an existing track, and whether the route is intended to accept one is not stated anywhere public, so the register records the capability and not a workflow.
- Audio sourceQ3 outputs speech, sound effects and music natively, on by default in the API, with a speech-only modethree tracks in one pass
- Audio sourceAudio is generated with the video rather than added afterwardsgenerated with the picture
- Audio sourceJoint audio and video generation, plus audio-to-video, on LTX-2.5two directions
- Audio sourceThe audio parameter defaults to true so the returned video carries an audio track, and false returns one withouton unless it is switched off
- What choosing silence costsEnabling or disabling audio does not affect pricing
- Audio sourceDescribed as a media generation model generating videos with synced audio, with video and audio both listed as outputgenerated with the picture
4Sources
Every line is read from the vendor documentation linked on the speech controls page, checked 2026-09-12. What counts as documented is on how read. Nearby: Which language, Two in frame, Who docs are for.