Sentioscope

Speech and voice controls, as each vendor documents them

Audio source: where the sound in a shot begins

Audio source is the first field these notes keep: whether speech leaves the model together with the picture, or the shot arrives silent and a voice is attached afterwards. Fifteen of the 17 entries fill it, which makes it the fullest column in the register. As of 2026-09-12.

Sound that arrives with the picture, against sound added laterThe two routes produce different schedules, different budgets and different failure modes. Every entry in this register documents the first; the second is what the field exists to rule out.Out of the same callAttached afterwardsStage that has to existNone; the take arrives finishedA dubbing pass, booked and scheduledLip movementAligned by constructionSomething to be matched, or patchedRedoing a shotA fresh performance comes with itThe picture changes, the voice can stayWhat the vendor statesWhere the sound originatesUsually nothing; a separate productSame footage, two production shapes
Fig. 1 The cell records which of these two a vendor describes, because the answer reshapes the production plan rather than one shot.
Controls documented, model by modelFilled where the model documents that control, hollow where nothing is published about it.Controls documented, model by modelWhere the sound comes fromWhere the soundcomes fromThe published wording behind the cellThe publishedwording behind th…D-IDD-ID — Where the sound comes from: From a script, or a supplied urlD-ID — The published wording behind the cell: A script is either text or an audio urlHedraHedra — Where the sound comes from: Supplied; audio is requiredHedra — The published wording behind the cell: Avatar videos are driven by audioHeyGenHeyGen — Where the sound comes from: From a speech endpoint, then renderedHeyGen — The published wording behind the cell: Text to speech turns a script into speech audioKling AIKling AI — Where the sound comes from: With the picture, on VIDEO 3.0Kling AI — The published wording behind the cell: Native audio in five languagesLTX StudioLTX Studio — Where the sound comes from: With the picture, and audio to videoLTX Studio — The published wording behind the cell: Joint audio and video generation on LTX-2.5Luma RayLuma Ray — Where the sound comes from: Not documented by the vendorLuma Ray — The published wording behind the cell: The video reference does not mention sound at allMiniMaxMiniMax — Where the sound comes from: With the pictureMiniMax — The published wording behind the cell: Native speech, lip movement following whoever is on screenPixVersePixVerse — Where the sound comes from: Supplied, or read from textPixVerse — The published wording behind the cell: A speech endpoint taking a clip or a synthesised voiceRunwayRunway — Where the sound comes from: Not documented by the vendorRunway — The published wording behind the cell: The voices page defines voices rather than shotsSceneMixerSceneMixer — Where the sound comes from: With the picture, in the language set for the projectSceneMixer — The published wording behind the cell: The line is performed while the shot rendersSora 2Sora 2 — Where the sound comes from: With the pictureSora 2 — The published wording behind the cell: Generating videos with synced audiosync-3sync-3 — Where the sound comes from: Supplied, or read from textsync-3 — The published wording behind the cell: Video or an image, paired with audio or with textSynthesiaSynthesia — Where the sound comes from: From the script, or uploadedSynthesia — The published wording behind the cell: Lip sync is generated from the spoken contentVeoVeo — Where the sound comes from: With the pictureVeo — The published wording behind the cell: Audio generated with the video rather than added afterwardsViduVidu — Where the sound comes from: With the picture, speech available on its ownVidu — The published wording behind the cell: Speech, effects and music in one pass, on by defaultWan 3.0Wan 3.0 — Where the sound comes from: With the picture, unless switched offWan 3.0 — The published wording behind the cell: The audio parameter defaults to trueWan2.2-S2VWan2.2-S2V — Where the sound comes from: Supplied; the model is audio-drivenWan2.2-S2V — The published wording behind the cell: Audio-driven cinematic video generation
Fig. 2 Filled where the model documents that control, hollow where nothing is published about it.
Models documenting each controlHow many models document each control. A hollow column is a statement about documentation, not capability.Models documenting each controlWhere the sound comes from15 of 17The published wording behind the ce…The published wording behind the cell15 of 17
Fig. 3 How many models document each control. A hollow column is a statement about documentation, not capability.
Where each vendor says the sound in a generated shot comes from. Recorded 2026-09-12.
ModelWhere the sound comes fromThe published wording behind the cell
D-IDFrom a script, or a supplied urlA script is either text or an audio url
HedraSupplied; audio is requiredAvatar videos are driven by audio
HeyGenFrom a speech endpoint, then renderedText to speech turns a script into speech audio
Kling AIWith the picture, on VIDEO 3.0Native audio in five languages
LTX StudioWith the picture, and audio to videoJoint audio and video generation on LTX-2.5
Luma RayNot documented by the vendorThe video reference does not mention sound at all
MiniMaxWith the pictureNative speech, lip movement following whoever is on screen
PixVerseSupplied, or read from textA speech endpoint taking a clip or a synthesised voice
RunwayNot documented by the vendorThe voices page defines voices rather than shots
SceneMixerWith the picture, in the language set for the projectThe line is performed while the shot renders
Sora 2With the pictureGenerating videos with synced audio
sync-3Supplied, or read from textVideo or an image, paired with audio or with text
SynthesiaFrom the script, or uploadedLip sync is generated from the spoken content
VeoWith the pictureAudio generated with the video rather than added afterwards
ViduWith the picture, speech available on its ownSpeech, effects and music in one pass, on by default
Wan 3.0With the picture, unless switched offThe audio parameter defaults to true
Wan2.2-S2VSupplied; the model is audio-drivenAudio-driven cinematic video generation

Inclusion rule. Models whose own documentation says where the sound in a generated shot originates. A separate voice product sold beside a silent generator does not earn a row in this field. Order. Alphabetical by model name.

1Why this field is read before the other three

Nothing else in a production plan survives a wrong answer here. Sound arriving with the picture means there is no dubbing stage, no separate voice budget and no synchronisation problem to solve. Silent footage means all three exist, and the schedule is a different shape from the first day of prep rather than from the first day of post.

So the cell holds a description of where the sound originates, not a yes or a no. Vidu's speech-only mode and the audio-to-video direction on LTX-2.5 sit in the same cell as Veo's flat sentence, because a reader is choosing between workflows rather than counting features.

2The field nobody leaves blank

Every model note here fills this cell, which makes it unlike the rest of the register. Voice source, per-character binding and the language list are each empty for somebody. Where the sound begins is not, because native audio is the claim a launch post can demonstrate in a clip and every vendor in this register has made it.

A full column is therefore less reassuring than it looks. A cell everybody fills is a cell everybody has a reason to fill, and the wording behind most of these is a product-page sentence rather than a parameter with limits attached to it.

3What a filled cell still leaves open

Reading the whole column settles where sound starts and nothing about whether it can be steered. Whether a chosen voice can be handed over is kept in voice source. Whether that voice survives into the next episode is kept in per-character binding. Which languages the speech is documented for is kept separately again, and a filled cell here predicts none of them.

Regeneration is the gap none of the fields close. Because the audio comes out of the same call as the picture, redoing a shot for a framing problem produces a fresh performance too, and nothing read for these notes documents a way to hold the take and replace the image.

4One entry at a time on this column

A cell gets a page of its own where the vendor says something specific in it, or where its silence is unusual among the entries answering the same way. The remaining cells are left in the table above, because a page repeating one short phrase would be worse than a row carrying it.

Entries that make the sound in the same pass as the picture:

  • Kling AI — with the picture, on video 3.0.
  • LTX Studio — with the picture, and audio to video.
  • MiniMax — with the picture.
  • SceneMixer — with the picture, in the language set for the project.
  • Sora 2 — with the picture.
  • Vidu — with the picture, speech available on its own.
  • Wan 3.0 — with the picture, unless switched off.

Entries that are handed a recording and fit a picture to it:

  • Hedra — supplied; audio is required.
  • sync-3 — supplied, or read from text.
  • Wan2.2-S2V — supplied; the model is audio-driven.

Entries that synthesise speech from a script as a step of its own:

  • D-ID — from a script, or a supplied url.
  • HeyGen — from a speech endpoint, then rendered.
  • PixVerse — supplied, or read from text.
  • Synthesia — from the script, or uploaded.

Entries that publish nothing about where the sound comes from:

  • Luma Ray — not documented by the vendor.
  • Audio source
    Native audio on VIDEO 3.0 in five languagesgenerated with the pictureKling AI, model guide / recorded 2026-09-12
  • Audio source
    Native speech with lip-sync tied to the speaker who is on screenspeaker-awareMiniMax, video generation guide / recorded 2026-09-12
  • Audio source
    Q3 outputs speech, sound effects and music natively, on by default in the API, with a speech-only modethree tracks in one passVidu, API reference / recorded 2026-09-12
  • Audio source
    Joint audio and video generation, plus audio-to-video, on LTX-2.5two directionsLTX Studio / recorded 2026-09-12
  • Audio source
    Audio is generated with the video rather than added afterwardsgenerated with the pictureGoogle, Veo documentation / recorded 2026-09-12
  • Audio source
    The video model is described as performing the line in the chosen language while it renders the shotgenerated with the pictureSceneMixer, languages guide / recorded 2026-09-12
  • Audio source
    Described as a media generation model generating videos with synced audio, with video and audio both listed as outputgenerated with the pictureOpenAI, Sora 2 model page / recorded 2026-09-22
  • Audio source
    The audio parameter defaults to true so the returned video carries an audio track, and false returns one withouton unless it is switched offAlibaba Cloud, Wan 3.0 API reference / recorded 2026-09-22
  • Audio source
    Avatar videos are driven by audio, and the character in the image will lip-sync and move to the audio providedhanded to the model rather than made by itHedra, avatar video guide / recorded 2026-09-22
  • Audio source
    Not documented by the vendor (as of 2026-09-22)no mention of sound in the video referenceLuma, video generation reference / recorded 2026-09-22
  • Audio source
    The accepted pairs are video with audio, video with text, image with audio and image with textsupplied or read from textSync, sync-3 model documentation / recorded 2026-09-22
  • Audio source
    A script is either text of up to forty thousand characters or an audio url, with audio limited to five minutes for clips and ten for talksa script read aloud or a fileD-ID, create a talk reference / recorded 2026-09-22
  • Audio source
    The card describes audio-driven cinematic video generation from an audio input with a reference image and an optional text prompthanded to the model rather than made by itWan-AI, Wan2.2-S2V-14B model card / recorded 2026-09-22
  • Audio source
    A text to speech endpoint turns a script into speech audio as a step of its ownan endpoint apart from the pictureHeyGen, API quick start / recorded 2026-09-22

5Sources

Each cell is read from the vendor page it links to, checked 2026-09-12. The fields sit side by side on the speech table, and what counts as documented is set out on how read. The other fields: Voice source, Per-character binding. All of them: the field notes.