Sentioscope

Speech and voice controls, as each vendor documents them

Entries that make the sound while they make the shot

Eight entries here produce sound in the same call as the picture. Alignment is a property of the model rather than of an input, so a shot can be attempted, heard and revised inside one pass, and the performance is whatever comes back. As of 2026-09-22.

One shared answer, then eight different second answersEvery entry on this route makes the sound with the picture and then diverges completely on what a production may hand over: a clip with a published ceiling, a preset library, a bound voice from nowhere stated, or nothing.Same architecture, and then what can you supplyA reference clipA published ceilingFifteen seconds, and theclip becomes an asset tostoreA library or a bound voiceA choice, or noneA voice exists; how it waschosen may not be statedNothing publishedWhatever comes backFour of the eight leavethis column entirely emptyThe first column is shared and the second decides the work
Fig. 1 Grouping by audio source explains a schedule and predicts almost nothing about whether a cast can be built.
Controls documented, model by modelFilled where the model documents that control, hollow where nothing is published about it.Controls documented, model by modelWhat the vendor says about the soundWhat the vendorsays about the…And what a production may hand overAnd what aproduction may…Kling AIKling AI — What the vendor says about the sound: With the picture, on VIDEO 3.0Kling AI — And what a production may hand over: Nothing documented as an inputLTX StudioLTX Studio — What the vendor says about the sound: With the picture, and audio to videoLTX Studio — And what a production may hand over: Nothing documented as an inputMiniMaxMiniMax — What the vendor says about the sound: With the pictureMiniMax — And what a production may hand over: A reference clip on the callSceneMixerSceneMixer — What the vendor says about the sound: With the picture, in the language set for the projectSceneMixer — And what a production may hand over: A preset library, or the user's own recordingsSora 2Sora 2 — What the vendor says about the sound: With the pictureSora 2 — And what a production may hand over: Nothing documented as an inputVeoVeo — What the vendor says about the sound: With the pictureVeo — And what a production may hand over: Not documented by the vendorViduVidu — What the vendor says about the sound: With the picture, speech available on its ownVidu — And what a production may hand over: Nothing documented as an inputWan 3.0Wan 3.0 — What the vendor says about the sound: With the picture, unless switched offWan 3.0 — And what a production may hand over: Reference audio, 15 seconds in total
Fig. 2 Filled where the model documents that control, hollow where nothing is published about it.
Sound made with the picture, beside what each of those vendors accepts as a voice. Recorded 2026-09-22.
ModelWhat the vendor says about the soundAnd what a production may hand over
Kling AIWith the picture, on VIDEO 3.0Nothing documented as an input
LTX StudioWith the picture, and audio to videoNothing documented as an input
MiniMaxWith the pictureA reference clip on the call
SceneMixerWith the picture, in the language set for the projectA preset library, or the user's own recordings
Sora 2With the pictureNothing documented as an input
VeoWith the pictureNot documented by the vendor
ViduWith the picture, speech available on its ownNothing documented as an input
Wan 3.0With the picture, unless switched offReference audio, 15 seconds in total

Inclusion rule. Entries whose own documentation says the sound leaves the model together with the picture. Products that are handed a recording, or that synthesise speech in a separate step, are listed on the other route pages. Order. Alphabetical by model name.

1Alignment is free, and direction is what you lose

Because the words and the frames come out of one call, nothing has to be matched to anything. That removes the failure every dubbing pipeline fights and replaces it with a different one: there is no separate performance to approve, so a line that comes back flat is a picture problem as well as an audio problem.

The cost lands on revision. Redoing a shot for a visual reason produces a fresh performance with it, and where a line was already approved that is an expense nobody planned. Productions working this way separate picture fixes from line fixes wherever the tool allows it, and most of these do not.

2The second column is where these eight stop looking alike

Every entry on this page answers the first question the same way and then diverges completely. Two accept a reference clip with a published ceiling in seconds. One names a preset library. Five publish nothing a production could supply, and two of those five bind a voice to a character anyway.

So the shared architecture predicts almost nothing about whether a cast can be built. Grouping by audio source is useful for understanding a schedule and useless for casting, which is why this register keeps the two as separate fields.

3What a silent competitor is not being compared with

Where audio is on by default, the baseline output of a model is a finished soundscape rather than footage waiting for post. A team benchmarking one of these against a silent generator is comparing two different deliverables, and the one that arrives with music will usually feel further along than it is.

Treating generated speech as a guide track that happens to be in sync, rather than as delivery audio, is the safer reading. Levels move between calls, room tone changes, and music comes and goes with the prompt.

4The entries on this route, one page each

Each of these links to that entry on the field this route groups by. The wording behind every cell, and the date it was read, sits on the page it links to.

  • Kling AI — with the picture, on video 3.0.
  • LTX Studio — with the picture, and audio to video.
  • MiniMax — with the picture.
  • SceneMixer — with the picture, in the language set for the project.
  • Sora 2 — with the picture.
  • Veo — with the picture.
  • Vidu — with the picture, speech available on its own.
  • Wan 3.0 — with the picture, unless switched off.
  • Audio source
    Native audio on VIDEO 3.0 in five languagesgenerated with the pictureKling AI, model guide / recorded 2026-09-12
  • Audio source
    Joint audio and video generation, plus audio-to-video, on LTX-2.5two directionsLTX Studio / recorded 2026-09-12
  • Audio source
    Audio is generated with the video rather than added afterwardsgenerated with the pictureGoogle, Veo documentation / recorded 2026-09-12
  • Audio source
    Q3 outputs speech, sound effects and music natively, on by default in the API, with a speech-only modethree tracks in one passVidu, API reference / recorded 2026-09-12
  • Audio source
    Described as a media generation model generating videos with synced audio, with video and audio both listed as outputgenerated with the pictureOpenAI, Sora 2 model page / recorded 2026-09-22
  • Audio source
    The audio parameter defaults to true so the returned video carries an audio track, and false returns one withouton unless it is switched offAlibaba Cloud, Wan 3.0 API reference / recorded 2026-09-22
  • Audio source
    The video model is described as performing the line in the chosen language while it renders the shotgenerated with the pictureSceneMixer, languages guide / recorded 2026-09-12
  • Audio source
    Native speech with lip-sync tied to the speaker who is on screenspeaker-awareMiniMax, video generation guide / recorded 2026-09-12

5Sources

Membership of this route is decided by the wording in the field it groups by, read from the documentation each vendor publishes. The column itself is on the field note, all five columns are on the speech table, and what counts as documented is on how read. Other routes: Handed a recording, Speech from a script.