Sentioscope

Speech and voice controls, as each vendor documents them

Vidu: three tracks in one pass, on by default

Q3 outputs speech, sound effects and music natively, on by default in the API, with a speech-only mode beside it. The default decides what most output from this model sounds like, because it applies to every call made without reading the parameter list. As of 2026-09-12.

Three tracks arrive by default, and one can arrive aloneQ3 outputs speech, sound effects and music natively and the API has them on by default. A speech-only mode keeps the dialogue and drops the rest, which is the mode a drama editor wants.What a default call returnsSpeechThe dialogue itself, the one track a drama edit needs to keep.Sound effectsGenerated with the shot, and baked into the returned file.MusicAlso generated, also baked in, and a creative decision an editorowns.Speech-only modeDocumented separately, and the reason this row readsdifferently.Silence and a clean stem both have to be asked for
Fig. 1 With the bed switched on by default, a scene meant to play dry will not unless the call says so, and the fix afterwards is a regeneration.
Vidu on audio source, statement by statement. Read from the vendor's API reference on 2026-09-12.
What the documentation settlesWhat it leaves to a take
Q3 outputs speech, sound effects and music natively, on by default in the APIHow the three tracks balance against each other in a returned file
A speech-only mode returns the dialogue without the effects and musicWhether that mode changes the delivery as well as the mix
Nothing is published about where a voice comes fromWhether a voice can be chosen, supplied or held to a character at all

Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.

1A default is a design decision disguised as a parameter

Audio on by default means the baseline output of this model is a finished soundscape rather than silent footage waiting for post. A team benchmarking it against a silent competitor is comparing two different deliverables, and the one that arrives with music will usually feel further along.

The same default makes silence something that has to be asked for. A scene meant to play dry will not unless the call says so, and the fix after the fact is a regeneration rather than a mute, because the bed is inside the file.

2Separating speech from the bed is the useful half

A speech-only mode is the one documented affordance in this register for keeping the dialogue and replacing everything under it. For drama that is the mode that matters, because the music is a creative decision that belongs to an editor rather than to a generator.

What the reference does not carry is any statement about the voice itself. Three tracks are described and nobody is identified as speaking, which is enough for a single clip and not enough for a cast.

3The others that make sound while they make the picture

Seven more entries generate audio in the same pass. This is the only one that documents a way to keep the speech and drop the rest.

  • Kling AI — with the picture, on video 3.0.
  • LTX Studio — with the picture, and audio to video.
  • MiniMax — with the picture.
  • SceneMixer — with the picture, in the language set for the project.
  • Sora 2 — with the picture.
  • Veo — with the picture.
  • Wan 3.0 — with the picture, unless switched off.
  • Audio source
    Q3 outputs speech, sound effects and music natively, on by default in the API, with a speech-only modethree tracks in one passVidu, API reference / recorded 2026-09-12
  • Per-character binding
    Not documented by the vendor (as of 2026-09-12)no public descriptionVidu, API reference / recorded 2026-09-12

4Sources

Read from the API reference at platform.vidu.com on 2026-09-12. The same column across every entry is on audio source; everything this vendor publishes about speech is on Vidu. What counts as documented is on how read.