Audio source: where the sound in a shot begins
Audio source is the first field these notes keep: whether speech leaves the model together with the picture, or the shot arrives silent and a voice is attached afterwards. Fifteen of the 17 entries fill it, which makes it the fullest column in the register. As of 2026-09-12.
| Model | Where the sound comes from | The published wording behind the cell |
|---|---|---|
| D-ID | From a script, or a supplied url | A script is either text or an audio url |
| Hedra | Supplied; audio is required | Avatar videos are driven by audio |
| HeyGen | From a speech endpoint, then rendered | Text to speech turns a script into speech audio |
| Kling AI | With the picture, on VIDEO 3.0 | Native audio in five languages |
| LTX Studio | With the picture, and audio to video | Joint audio and video generation on LTX-2.5 |
| Luma Ray | Not documented by the vendor | The video reference does not mention sound at all |
| MiniMax | With the picture | Native speech, lip movement following whoever is on screen |
| PixVerse | Supplied, or read from text | A speech endpoint taking a clip or a synthesised voice |
| Runway | Not documented by the vendor | The voices page defines voices rather than shots |
| SceneMixer | With the picture, in the language set for the project | The line is performed while the shot renders |
| Sora 2 | With the picture | Generating videos with synced audio |
| sync-3 | Supplied, or read from text | Video or an image, paired with audio or with text |
| Synthesia | From the script, or uploaded | Lip sync is generated from the spoken content |
| Veo | With the picture | Audio generated with the video rather than added afterwards |
| Vidu | With the picture, speech available on its own | Speech, effects and music in one pass, on by default |
| Wan 3.0 | With the picture, unless switched off | The audio parameter defaults to true |
| Wan2.2-S2V | Supplied; the model is audio-driven | Audio-driven cinematic video generation |
Inclusion rule. Models whose own documentation says where the sound in a generated shot originates. A separate voice product sold beside a silent generator does not earn a row in this field. Order. Alphabetical by model name.
1Why this field is read before the other three
Nothing else in a production plan survives a wrong answer here. Sound arriving with the picture means there is no dubbing stage, no separate voice budget and no synchronisation problem to solve. Silent footage means all three exist, and the schedule is a different shape from the first day of prep rather than from the first day of post.
So the cell holds a description of where the sound originates, not a yes or a no. Vidu's speech-only mode and the audio-to-video direction on LTX-2.5 sit in the same cell as Veo's flat sentence, because a reader is choosing between workflows rather than counting features.
2The field nobody leaves blank
Every model note here fills this cell, which makes it unlike the rest of the register. Voice source, per-character binding and the language list are each empty for somebody. Where the sound begins is not, because native audio is the claim a launch post can demonstrate in a clip and every vendor in this register has made it.
A full column is therefore less reassuring than it looks. A cell everybody fills is a cell everybody has a reason to fill, and the wording behind most of these is a product-page sentence rather than a parameter with limits attached to it.
3What a filled cell still leaves open
Reading the whole column settles where sound starts and nothing about whether it can be steered. Whether a chosen voice can be handed over is kept in voice source. Whether that voice survives into the next episode is kept in per-character binding. Which languages the speech is documented for is kept separately again, and a filled cell here predicts none of them.
Regeneration is the gap none of the fields close. Because the audio comes out of the same call as the picture, redoing a shot for a framing problem produces a fresh performance too, and nothing read for these notes documents a way to hold the take and replace the image.
4One entry at a time on this column
A cell gets a page of its own where the vendor says something specific in it, or where its silence is unusual among the entries answering the same way. The remaining cells are left in the table above, because a page repeating one short phrase would be worse than a row carrying it.
Entries that make the sound in the same pass as the picture:
- Kling AI — with the picture, on video 3.0.
- LTX Studio — with the picture, and audio to video.
- MiniMax — with the picture.
- SceneMixer — with the picture, in the language set for the project.
- Sora 2 — with the picture.
- Vidu — with the picture, speech available on its own.
- Wan 3.0 — with the picture, unless switched off.
Entries that are handed a recording and fit a picture to it:
- Hedra — supplied; audio is required.
- sync-3 — supplied, or read from text.
- Wan2.2-S2V — supplied; the model is audio-driven.
Entries that synthesise speech from a script as a step of its own:
- D-ID — from a script, or a supplied url.
- HeyGen — from a speech endpoint, then rendered.
- PixVerse — supplied, or read from text.
- Synthesia — from the script, or uploaded.
Entries that publish nothing about where the sound comes from:
- Luma Ray — not documented by the vendor.
- Audio sourceNative audio on VIDEO 3.0 in five languagesgenerated with the picture
- Audio sourceNative speech with lip-sync tied to the speaker who is on screenspeaker-aware
- Audio sourceQ3 outputs speech, sound effects and music natively, on by default in the API, with a speech-only modethree tracks in one pass
- Audio sourceJoint audio and video generation, plus audio-to-video, on LTX-2.5two directions
- Audio sourceAudio is generated with the video rather than added afterwardsgenerated with the picture
- Audio sourceThe video model is described as performing the line in the chosen language while it renders the shotgenerated with the picture
- Audio sourceDescribed as a media generation model generating videos with synced audio, with video and audio both listed as outputgenerated with the picture
- Audio sourceThe audio parameter defaults to true so the returned video carries an audio track, and false returns one withouton unless it is switched off
- Audio sourceAvatar videos are driven by audio, and the character in the image will lip-sync and move to the audio providedhanded to the model rather than made by it
- Audio sourceNot documented by the vendor (as of 2026-09-22)no mention of sound in the video reference
- Audio sourceThe accepted pairs are video with audio, video with text, image with audio and image with textsupplied or read from text
- Audio sourceA script is either text of up to forty thousand characters or an audio url, with audio limited to five minutes for clips and ten for talksa script read aloud or a file
- Audio sourceThe card describes audio-driven cinematic video generation from an audio input with a reference image and an optional text prompthanded to the model rather than made by it
- Audio sourceA text to speech endpoint turns a script into speech audio as a step of its ownan endpoint apart from the picture
5Sources
Each cell is read from the vendor page it links to, checked 2026-09-12. The fields sit side by side on the speech table, and what counts as documented is set out on how read. The other fields: Voice source, Per-character binding. All of them: the field notes.