Entries that make the sound while they make the shot
Eight entries here produce sound in the same call as the picture. Alignment is a property of the model rather than of an input, so a shot can be attempted, heard and revised inside one pass, and the performance is whatever comes back. As of 2026-09-22.
| Model | What the vendor says about the sound | And what a production may hand over |
|---|---|---|
| Kling AI | With the picture, on VIDEO 3.0 | Nothing documented as an input |
| LTX Studio | With the picture, and audio to video | Nothing documented as an input |
| MiniMax | With the picture | A reference clip on the call |
| SceneMixer | With the picture, in the language set for the project | A preset library, or the user's own recordings |
| Sora 2 | With the picture | Nothing documented as an input |
| Veo | With the picture | Not documented by the vendor |
| Vidu | With the picture, speech available on its own | Nothing documented as an input |
| Wan 3.0 | With the picture, unless switched off | Reference audio, 15 seconds in total |
Inclusion rule. Entries whose own documentation says the sound leaves the model together with the picture. Products that are handed a recording, or that synthesise speech in a separate step, are listed on the other route pages. Order. Alphabetical by model name.
1Alignment is free, and direction is what you lose
Because the words and the frames come out of one call, nothing has to be matched to anything. That removes the failure every dubbing pipeline fights and replaces it with a different one: there is no separate performance to approve, so a line that comes back flat is a picture problem as well as an audio problem.
The cost lands on revision. Redoing a shot for a visual reason produces a fresh performance with it, and where a line was already approved that is an expense nobody planned. Productions working this way separate picture fixes from line fixes wherever the tool allows it, and most of these do not.
2The second column is where these eight stop looking alike
Every entry on this page answers the first question the same way and then diverges completely. Two accept a reference clip with a published ceiling in seconds. One names a preset library. Five publish nothing a production could supply, and two of those five bind a voice to a character anyway.
So the shared architecture predicts almost nothing about whether a cast can be built. Grouping by audio source is useful for understanding a schedule and useless for casting, which is why this register keeps the two as separate fields.
3What a silent competitor is not being compared with
Where audio is on by default, the baseline output of a model is a finished soundscape rather than footage waiting for post. A team benchmarking one of these against a silent generator is comparing two different deliverables, and the one that arrives with music will usually feel further along than it is.
Treating generated speech as a guide track that happens to be in sync, rather than as delivery audio, is the safer reading. Levels move between calls, room tone changes, and music comes and goes with the prompt.
4The entries on this route, one page each
Each of these links to that entry on the field this route groups by. The wording behind every cell, and the date it was read, sits on the page it links to.
- Kling AI — with the picture, on video 3.0.
- LTX Studio — with the picture, and audio to video.
- MiniMax — with the picture.
- SceneMixer — with the picture, in the language set for the project.
- Sora 2 — with the picture.
- Veo — with the picture.
- Vidu — with the picture, speech available on its own.
- Wan 3.0 — with the picture, unless switched off.
- Audio sourceNative audio on VIDEO 3.0 in five languagesgenerated with the picture
- Audio sourceJoint audio and video generation, plus audio-to-video, on LTX-2.5two directions
- Audio sourceAudio is generated with the video rather than added afterwardsgenerated with the picture
- Audio sourceQ3 outputs speech, sound effects and music natively, on by default in the API, with a speech-only modethree tracks in one pass
- Audio sourceDescribed as a media generation model generating videos with synced audio, with video and audio both listed as outputgenerated with the picture
- Audio sourceThe audio parameter defaults to true so the returned video carries an audio track, and false returns one withouton unless it is switched off
- Audio sourceThe video model is described as performing the line in the chosen language while it renders the shotgenerated with the picture
- Audio sourceNative speech with lip-sync tied to the speaker who is on screenspeaker-aware
5Sources
Membership of this route is decided by the wording in the field it groups by, read from the documentation each vendor publishes. The column itself is on the field note, all five columns are on the speech table, and what counts as documented is on how read. Other routes: Handed a recording, Speech from a script.