# Sentioscope (sentioscope.com) > Sentioscope records what each AI video model documents about generated speech: whether audio comes out of the model, how a voice is bound to a character, and what the vendor says about lip-sync. This site is open for crawling, indexing, quoting and model training. Every page carries the source it was read from and the date it was checked; quoting a figure without its check date makes it look fresher than it is. Full text of every page in one file: https://sentioscope.com/llms-full.txt ## Overview - [Sentioscope: speech controls in video models](https://sentioscope.com/): What AI video models document about generated speech: audio source, languages, how a voice attaches to a character, and lip-sync. ## Control table - [Where the voice comes from, model by model](https://sentioscope.com/controls/speech/): Seventeen models on four speech fields: where audio comes from, which languages are named, how a voice attaches to a character, and what drives lip-sync. ## Models - [What D-ID documents about speech](https://sentioscope.com/models/d-id/dialogue/): What D-ID documents about generated speech: where the audio comes from, how a voice attaches to a character, and whether lip-sync is covered. - [What Hedra documents about speech](https://sentioscope.com/models/hedra/dialogue/): What Hedra documents about generated speech: where the audio comes from, how a voice attaches to a character, and whether lip-sync is covered. - [What HeyGen documents about speech](https://sentioscope.com/models/heygen/dialogue/): What HeyGen documents about generated speech: where the audio comes from, how a voice attaches to a character, and whether lip-sync is covered. - [What Kling AI documents about speech](https://sentioscope.com/models/kling-ai/dialogue/): What Kling AI documents about generated speech: where the audio comes from, how a voice attaches to a character, and whether lip-sync is covered. - [What LTX Studio documents about speech](https://sentioscope.com/models/ltx-studio/dialogue/): What LTX Studio documents about generated speech: where the audio comes from, how a voice attaches to a character, and whether lip-sync is covered. - [What Luma Ray documents about speech](https://sentioscope.com/models/luma-ray/dialogue/): What Luma Ray documents about generated speech: where the audio comes from, how a voice attaches to a character, and whether lip-sync is covered. - [What MiniMax documents about speech](https://sentioscope.com/models/minimax/dialogue/): What MiniMax documents about generated speech: where the audio comes from, how a voice attaches to a character, and whether lip-sync is covered. - [What PixVerse documents about speech](https://sentioscope.com/models/pixverse/dialogue/): What PixVerse documents about generated speech: where the audio comes from, how a voice attaches to a character, and whether lip-sync is covered. - [What Runway documents about speech](https://sentioscope.com/models/runway/dialogue/): What Runway documents about generated speech: where the audio comes from, how a voice attaches to a character, and whether lip-sync is covered. - [What SceneMixer documents about speech](https://sentioscope.com/models/scenemixer/dialogue/): What SceneMixer documents about generated speech: where the audio comes from, how a voice attaches to a character, and whether lip-sync is covered. - [What Sora 2 documents about speech](https://sentioscope.com/models/sora-2/dialogue/): What Sora 2 documents about generated speech: where the audio comes from, how a voice attaches to a character, and whether lip-sync is covered. - [What sync-3 documents about speech](https://sentioscope.com/models/sync-3/dialogue/): What sync-3 documents about generated speech: where the audio comes from, how a voice attaches to a character, and whether lip-sync is covered. - [What Synthesia documents about speech](https://sentioscope.com/models/synthesia/dialogue/): What Synthesia documents about generated speech: where the audio comes from, how a voice attaches to a character, and whether lip-sync is covered. - [What Veo documents about speech](https://sentioscope.com/models/veo/dialogue/): What Veo documents about generated speech: where the audio comes from, how a voice attaches to a character, and whether lip-sync is covered. - [What Vidu documents about speech](https://sentioscope.com/models/vidu/dialogue/): What Vidu documents about generated speech: where the audio comes from, how a voice attaches to a character, and whether lip-sync is covered. - [What Wan 3.0 documents about speech](https://sentioscope.com/models/wan-3/dialogue/): What Wan 3.0 documents about generated speech: where the audio comes from, how a voice attaches to a character, and whether lip-sync is covered. - [What Wan2.2-S2V documents about speech](https://sentioscope.com/models/wan-s2v/dialogue/): What Wan2.2-S2V documents about generated speech: where the audio comes from, how a voice attaches to a character, and whether lip-sync is covered. ## How read - [How a field here gets filled in](https://sentioscope.com/reading/): The five speech fields recorded for every model here, why no cell is ever filled in from generated output, and what a not-documented cell actually means. ## Questions - [How a character keeps the same voice](https://sentioscope.com/questions/keeping-a-voice/): Four models keep a voice as a stored object you approve once. Five rebuild it on every call. The rest publish nothing at all about persistence. - [Lip-sync answered four different ways](https://sentioscope.com/questions/lip-sync-three-answers/): One vendor describes a mechanism, one names the feature without describing it, one removes the step, and several take the audio in so sync is the task. - [Where the voices actually come from](https://sentioscope.com/questions/where-voices-come-from/): A catalogue, a clone from a recording, a finished track, or a voice described in words. One route is ruled out on likeness grounds. - [Which models return sound unless told otherwise](https://sentioscope.com/questions/audio-on-by-default/): Two vendors state that sound comes back unless it is switched off, and one of them publishes what switching it off does to the price. - [Which language the performance is in](https://sentioscope.com/questions/language-of-the-performance/): One vendor publishes a catalogue with language codes, one names fifteen dialogue languages, six publish a count or a claim, and nine publish nothing. - [What happens with two speakers in frame](https://sentioscope.com/questions/two-speakers/): Two characters in frame with a line each is the commonest shot in drama, and exactly one vendor publishes any guidance for it. - [Why API documentation answers the wrong questions](https://sentioscope.com/questions/what-the-docs-are-for/): Developer references describe parameters. A serial production needs behaviour, and the two audiences rarely end up reading the same page. - [How long a generated shot with speech can run](https://sentioscope.com/questions/how-long-a-shot-can-be/): Eight entries publish a length for a generation or its audio, from a four-second shot to a ten-minute maximum. The rest publish nothing. - [What fifteen seconds of reference audio can carry](https://sentioscope.com/questions/fifteen-seconds-of-voice/): Two entries cap reference audio at fifteen seconds in total, and the published windows elsewhere run from one recording to twenty minutes of studio time. - [Which answers here can be checked by software](https://sentioscope.com/questions/machine-checkable-answers/): Most of this register is prose. Four entries publish something a script could verify: a language code, a stable id, an accepted format or a byte ceiling. - [What it costs to ask for a shot with no sound](https://sentioscope.com/questions/what-silence-costs/): One entry states plainly that turning audio off does not change the rate. Two fold audio into a per-second price, and thirteen say nothing at all. - [Where a spoken line goes in a prompt](https://sentioscope.com/questions/where-a-line-goes/): Four entries tell a writer where dialogue belongs: a labelled block, a clause after one keyword, a quoted line, or a reference voice instead of words. - [Whether a voice can be heard before a shot exists](https://sentioscope.com/questions/can-a-voice-be-auditioned/): One entry documents a preview on a voice that has reached a ready state. No other entry describes any way to hear a voice before generating with it. - [Which entries name somebody else as the voice supplier](https://sentioscope.com/questions/outside-suppliers-named/): Two entries name outside companies for speech: one lists five providers behind its voice field, and one names an audio house as a planned integration. - [Which column every entry manages to fill](https://sentioscope.com/questions/the-field-that-is-never-blank/): Audio source is filled by fifteen of the seventeen entries. Every other column in this register is blank for most of them, and lip-sync is blank for seven. - [Whether the dialogue can be kept without the music](https://sentioscope.com/questions/speech-without-the-bed/): One entry documents a speech-only mode. Everywhere else generated audio arrives as one file, so replacing the music means regenerating the shot. - [What happens to the voice when a shot is redone](https://sentioscope.com/questions/regenerating-a-shot/): Where audio comes out with the picture, redoing a shot produces a new performance. Where it was supplied, the approved read survives the re-render. - [Whether any vendor says whose voice may be used](https://sentioscope.com/questions/consent-in-the-documentation/): One entry treats cloning an identifiable voice as a likeness question and names digital replica statutes. The rest document cloning and skip permission. - [Entries that document a second lever beside the voice](https://sentioscope.com/questions/a-second-control/): Four entries publish a control most of the register lacks: a pose track, a speech-only mode, a reverse audio route, or a voice asked for in words. - [What kind of document each answer was read from](https://sentioscope.com/questions/what-kind-of-page-it-was/): The shape of an answer follows the kind of page it sits on: an API reference, a model card, a prompting guide or a compliance checklist. - [Which containers and file sizes are accepted](https://sentioscope.com/questions/formats-and-file-sizes/): Four entries publish what they accept as a file: WAV or MP3, a ten-megabyte ceiling, or a hundred megabytes on both sides of one call. - [What decides how long a generated shot runs](https://sentioscope.com/questions/what-decides-a-shot-length/): Length comes from three places: a menu of fixed options, a range a caller chooses, or the recording a production supplies. - [Whether generated audio is charged for separately](https://sentioscope.com/questions/who-charges-for-audio/): No entry here bills audio as a separate line. Two fold it into a per-second rate, and one says that switching it off changes nothing. ## Explainers - [Writing dialogue a model can perform](https://sentioscope.com/learn/writing-lines-for-generation/): Generated speech follows the text more literally than an actor does. Punctuation, numbers and abbreviations all become choices the writer is making. - [Dubbing against regenerating a language](https://sentioscope.com/learn/dubbing-or-regenerating/): A second language can be reached by replacing the audio or by generating again. They produce different problems and cost in different places. - [How lip-sync is produced, and where it breaks](https://sentioscope.com/learn/how-lip-sync-works/): Generated speech can be produced with the picture, laid over it afterwards, or driven from a supplied track. Each route fails in a different place. - [Casting a voice you can keep for forty episodes](https://sentioscope.com/learn/casting-a-generated-voice/): Choosing a generated voice is a casting decision with a constraint attached: whatever you pick has to be reproducible months later, by somebody else. - [Three ways to get a voice, and what each costs](https://sentioscope.com/learn/three-routes-to-a-voice/): A preset library, a cloned voice, or speech generated with the picture. They differ in price, in what they need from you, and in lip alignment. - [How to read a cell that records nothing](https://sentioscope.com/learn/reading-a-blank-cell/): An empty cell says a vendor has not published something. It does not say the product cannot do it, and the difference decides what to do next. - [Why a number of languages settles nothing](https://sentioscope.com/learn/counts-are-not-lists/): A total says a capability is broad. A list says whether the market a production ships to next quarter is inside it, which is the real question. - [Covering a conversation when nothing promises the mouth](https://sentioscope.com/learn/two-speakers-in-one-shot/): The commonest shot in drama is two people with one of them talking, and almost nothing says which face will move. Coverage is the lever that remains. - [Filing the clip that made a character sound right](https://sentioscope.com/learn/keeping-a-clip-with-a-character/): Where a voice is rebuilt from a reference clip on every call, the archive is what holds a cast together, and the output never records the clip. - [Scheduling around audio that has to exist first](https://sentioscope.com/learn/scheduling-a-supplied-track/): Where a recording drives the generation, casting and a session sit in front of the first frame, which reshapes the whole schedule. - [Why a blank on a model card means something else](https://sentioscope.com/learn/what-open-weights-change/): On a hosted product an absent sentence is a refusal to commit. Where the weights are published, it is a measurement nobody has taken. - [Separating a picture revision from an audio one](https://sentioscope.com/learn/picture-fixes-and-line-fixes/): Where sound and picture come from one call, redoing a shot for a visual reason returns a new performance and an approved line disappears without warning. - [How framing decides whether alignment holds up](https://sentioscope.com/learn/framing-and-the-mouth/): One published condition says lip-sync performs better with a figure framed closer to camera, which turns a mouth problem into a coverage decision. - [Hearing a voice before a season depends on it](https://sentioscope.com/learn/auditioning-before-a-season/): One entry documents a preview. Everywhere else the first generation is the audition, so a short deliberate test beforehand is the substitute. - [Sung lines and read copy are not dialogue](https://sentioscope.com/learn/singing-and-read-copy/): One vendor names singing and advertisements as supported audio types. Both have timing and breath requirements that dialogue does not. - [What to find out before a platform gets chosen](https://sentioscope.com/learn/free-tiers-and-first-tests/): The questions that decide a series are the ones documentation answers least, and most of them can be settled with a handful of short generations. - [A line that fits a shot can still be unreadable](https://sentioscope.com/learn/subtitles-and-spoken-length/): Spoken length and subtitle reading speed are two separate ceilings, and a muted viewer gets the subtitle as the whole performance. - [The sentence that would settle each column](https://sentioscope.com/learn/asking-the-right-question/): Where documentation is silent, it is worth knowing exactly what a vendor would have to write for the question to become answerable at all. - [Why a quiet scene has to be asked for explicitly](https://sentioscope.com/learn/when-silence-is-asked-for/): Where audio is on by default, a scene meant to play dry will not, and the fix is a regeneration rather than a mute: the bed is in the file. - [Doing the arithmetic before a call is spent](https://sentioscope.com/learn/sizing-a-scene/): A shot with a published ceiling holds a predictable number of words, and the sum takes less time than regenerating a line that overruns. - [The document that holds a generated cast together](https://sentioscope.com/learn/a-cast-list-for-dialogue/): Only two entries attach a voice to a character object, and even there nothing says how to rebuild one, so the cast list stays a production document. ## Terms - [Audio generated with the picture, not after it](https://sentioscope.com/terms/native-audio/): Four vendors make speech in the same pass as the shot and publish a clip length with it. The switch, the price and the dialogue syntax differ on all four. - [Supplying a voice instead of describing one](https://sentioscope.com/terms/reference-audio/): Four vendors accept a recording as the source of a voice, with caps on how many and how long. The limits are tight enough to shape casting decisions. - [How many words fit in a generated shot](https://sentioscope.com/terms/dialogue-timing/): Speech runs at a measurable rate and shots have fixed ceilings. The arithmetic between them decides how a line has to be written. - [Lip-sync: matching a mouth to a waveform](https://sentioscope.com/terms/lip-sync/): Lip-sync covers three different mechanisms with three different failures, and most vendors name the word without saying which of the three applies. - [Voice id: a handle that outlives one call](https://sentioscope.com/terms/voice-id/): A voice id is a short stable reference to a voice, and it is the cheapest continuity mechanism in this register because it cannot drift. - [Voice cloning: building a voice from a recording](https://sentioscope.com/terms/voice-cloning/): Cloning turns a recording into a reusable voice, and the published thresholds run from one file to twenty minutes or more of studio time. - [Text to speech: making the audio a stage of its own](https://sentioscope.com/terms/text-to-speech/): Where speech is synthesised in a separate call, the audio becomes an artefact that can be fetched, approved and reused before any frame exists. - [Audio-driven video: the recording as the instruction](https://sentioscope.com/terms/audio-driven-video/): Three entries treat a supplied waveform as the driving input, which puts casting and a session in front of the first frame and names the driver for free. - [Speaker label: naming who talks inside a prompt](https://sentioscope.com/terms/speaker-label/): A speaker label routes a line to a face within one generation, and carries no identity into the next, which is the distinction most often read past. - [Voice catalogue: an inventory you can choose from](https://sentioscope.com/terms/voice-catalogue/): A catalogue is a published set of voices to pick from, and the gap between a browsable one and an asserted one decides whether casting is possible. - [Language code: a language written for software](https://sentioscope.com/terms/language-code/): A language name can mean a market, a script or a dialect. A code picks one, which is why a catalogue with codes outranks a total with none. - [Guide track: audio in sync but not for delivery](https://sentioscope.com/terms/guide-track/): Generated speech arrives aligned by construction, which is the hard part, and is usually not a finished mix, which a post budget has to assume. - [Dubbing pass: a new performance over old frames](https://sentioscope.com/terms/dubbing-pass/): A dubbing pass records a translated performance and lays it over finished footage, which is the stage a lip-sync repair exists to clean up. - [Timbre: the part of a voice a short clip carries](https://sentioscope.com/terms/timbre/): Timbre is the colour of a voice, and it is the property that transfers reliably from a very short reference clip. Accent and pacing do not. - [Digital replica: a voice treated as a likeness](https://sentioscope.com/terms/digital-replica/): One entry in this register treats cloning an identifiable voice as a likeness question and names United States digital replica statutes beside it. - [Clip length: the unit a line has to fit inside](https://sentioscope.com/terms/clip-length/): A generated shot has a published ceiling, and that ceiling decides how a line has to be written long before anybody spends a call on it. - [On-screen speaker: which face the sound belongs to](https://sentioscope.com/terms/on-screen-speaker/): With two people in frame, something has to decide whose mouth moves. Four entries address it, and no two of them address it the same way. ## Fields - [Audio source: where the sound in a shot begins](https://sentioscope.com/fields/audio-source/): Audio source records whether a model returns speech in the same pass as the picture or leaves the shot silent, with the published wording behind each cell. - [Voice source: what a production hands over](https://sentioscope.com/fields/voice-source/): Voice source records what a production can supply so speech follows a chosen voice: a preset library, a short reference clip, or nothing published. - [Per-character binding: does a voice persist](https://sentioscope.com/fields/per-character-binding/): Per-character binding records whether a voice belongs to a reusable character or is supplied again on every call, and what that costs across a series. - [Languages: which ones are named for dialogue](https://sentioscope.com/fields/languages/): Languages records which spoken languages a vendor names for dialogue, and whether the documentation gives a list to plan against or only a count. - [Lip-sync: what the vendor says moves the mouth](https://sentioscope.com/fields/lip-sync/): Lip-sync records whether a vendor names something that drives the mouth, claims the feature without a driver, or leaves the question off the page. ## One model, one field - [D-ID on audio source: a script, or a supplied file](https://sentioscope.com/fields/audio-source/d-id/): D-ID takes a text script of up to forty thousand characters or an audio url, with ceilings of five minutes for a clip and ten for a talk. - [Hedra on audio source: the recording comes first](https://sentioscope.com/fields/audio-source/hedra/): Hedra says avatar videos are driven by audio and that the audio normally decides the video length, so the take is settled before a frame exists. - [HeyGen on audio source: speech from its own endpoint](https://sentioscope.com/fields/audio-source/heygen/): HeyGen documents a text to speech endpoint that turns a script into speech audio as a step of its own, before any picture is rendered against it. - [Kling AI on audio source: native, on one model line](https://sentioscope.com/fields/audio-source/kling-ai/): Kling AI documents native audio on VIDEO 3.0 in five languages, which attaches the capability to a model version rather than to a plan or an account. - [LTX Studio on audio source: a route each way](https://sentioscope.com/fields/audio-source/ltx-studio/): LTX Studio documents joint audio and video generation on LTX-2.5 and an audio-to-video route where an existing track drives the picture instead. - [Luma Ray on audio source: the reference is silent](https://sentioscope.com/fields/audio-source/luma-ray/): Luma's video generation reference names models and resolutions and never mentions sound, while a separate page lists audio integrations as a plan. - [MiniMax on audio source: speech, and who says it](https://sentioscope.com/fields/audio-source/minimax/): MiniMax documents native speech generated with the picture, with lip movement tied to whoever is on screen, which implies a choice the prompt cannot see. - [PixVerse on audio source: a speech endpoint first](https://sentioscope.com/fields/audio-source/pixverse/): PixVerse routes speech through an endpoint taking a clip or a synthesised voice, with audio and video each capped at sixty seconds and a hundred megabytes. - [SceneMixer on audio source: performed, not dubbed](https://sentioscope.com/fields/audio-source/scenemixer/): SceneMixer states the model performs the line in the language set for the project while it renders the shot, so no track is laid over the footage. - [Sora 2 on audio source: audio, and a clip length](https://sentioscope.com/fields/audio-source/sora-2/): Sora 2 is described as generating videos with synced audio, and the guide publishes clip lengths of sixteen or twenty seconds, extendable to 120. - [sync-3 on audio source: four accepted input pairs](https://sentioscope.com/fields/audio-source/sync-3/): sync-3 accepts video with audio, video with text, image with audio and image with text, so the sound is supplied or read from a script on the call. - [Synthesia on audio source: the script, or an upload](https://sentioscope.com/fields/audio-source/synthesia/): Synthesia generates lip sync and facial expressions from the spoken content, so the script is the source of both the audio and the performance on screen. - [Vidu on audio source: three tracks, on by default](https://sentioscope.com/fields/audio-source/vidu/): Vidu documents speech, sound effects and music generated natively on Q3, switched on by default in the API, with a mode that returns speech alone. - [Wan 3.0 on audio source: on unless switched off](https://sentioscope.com/fields/audio-source/wan-3/): Wan 3.0 defaults its audio parameter to true, and the reference states that enabling or disabling audio does not affect pricing at all. - [Wan2.2-S2V on audio source: audio drives the shot](https://sentioscope.com/fields/audio-source/wan-s2v/): The Wan2.2-S2V card describes audio-driven cinematic video generation from an audio input with a reference image, and publishes weights under Apache 2.0. - [D-ID on languages: a field, and nothing to fill it](https://sentioscope.com/fields/languages/d-id/): D-ID's talk reference puts an optional language beside the chosen voice, and publishes no list of the values that field will accept. - [Hedra on languages: a claim with nothing named](https://sentioscope.com/fields/languages/hedra/): Hedra's model page claims full multi-language support and a ten-minute maximum duration, and names none of the languages that claim covers. - [HeyGen on languages: a count from translation](https://sentioscope.com/fields/languages/heygen/): HeyGen documents thirty or more languages for video translation, with cloning and lip-sync named beside them, and no list attached to the count. - [Kling AI on languages: five, none of them named](https://sentioscope.com/fields/languages/kling-ai/): Kling AI counts five languages for native audio on VIDEO 3.0 and names none of them, which leaves the question a commissioning meeting asks unanswered. - [PixVerse on languages: multiple, and three reads](https://sentioscope.com/fields/languages/pixverse/): PixVerse says multiple languages and audio types are supported, including speech, singing and advertisements, and names no language among them. - [Runway on languages: multilingual, nothing named](https://sentioscope.com/fields/languages/runway/): Runway's custom voices reference offers a multilingual model option and names no language, so this column stays empty on the page that defines voices. - [SceneMixer on languages: fifteen named, plus one](https://sentioscope.com/fields/languages/scenemixer/): SceneMixer names fifteen dialogue languages set per project, with Cantonese a sixteenth for dialogue only, which makes it one of two lists in this column. - [Sora 2 on languages: dialogue advice, no language](https://sentioscope.com/fields/languages/sora-2/): Sora 2's prompting guide says where a line of dialogue goes and how to label two speakers, and never mentions which language the line is in. - [sync-3 on languages: ninety-five, counted not named](https://sentioscope.com/fields/languages/sync-3/): sync-3 supports ninety-five or more languages, described as the same coverage as the models before it, and that total arrives with no list. - [Synthesia on languages: a catalogue, with codes](https://sentioscope.com/fields/languages/synthesia/): Synthesia lists every voice with a formal and native language name, a language code, a gender, a name and an id, which is a catalogue rather than a count. - [Wan 3.0 on languages: caps in writing, no language](https://sentioscope.com/fields/languages/wan-3/): Wan 3.0 publishes audio defaults, reference-audio caps in seconds and megabytes and a dialogue syntax, and never names a spoken language. - [Wan2.2-S2V on languages: whatever was recorded](https://sentioscope.com/fields/languages/wan-s2v/): The Wan2.2-S2V card publishes nothing about spoken language, which follows from an audio-driven model: the language is whatever the recording holds. - [D-ID on voice source: five providers, named](https://sentioscope.com/fields/voice-source/d-id/): D-ID names five speech providers for the voice, from Microsoft to Azure OpenAI, and the voice itself is a voice id chosen from their lists. - [Hedra on voice source: a track, or a voice id](https://sentioscope.com/fields/voice-source/hedra/): Hedra takes an uploaded audio track, or generates speech inline by naming a voice id drawn from the voices endpoint, so both routes are documented. - [HeyGen on voice source: two grades of clone](https://sentioscope.com/fields/voice-source/heygen/): HeyGen documents a voice clone that is instant from one recording or professional from twenty minutes or more, then passed as a voice id. - [Kling AI on voice source: nothing to hand over](https://sentioscope.com/fields/voice-source/kling-ai/): Kling AI binds a voice to an element so a character keeps it, and publishes nothing a production could supply to decide what that voice is. - [LTX Studio on voice source: attached, not supplied](https://sentioscope.com/fields/voice-source/ltx-studio/): LTX Studio attaches voices to character Elements and never says where those voices come from, so nothing is documented as an input on this field. - [MiniMax on voice source: fifteen seconds in total](https://sentioscope.com/fields/voice-source/minimax/): MiniMax caps reference audio at fifteen seconds in total across at most three clips, which is the tightest published figure in this column. - [PixVerse on voice source: built-in, or from a sample](https://sentioscope.com/fields/voice-source/pixverse/): PixVerse documents both built-in voices and custom voices created from user-provided sample audio, with the chosen speaker passed as an id. - [Runway on voice source: a sample, or a sentence](https://sentioscope.com/fields/voice-source/runway/): Runway builds a custom voice from a sample of ten seconds to five minutes and at most 10 MB, or from a written description of at least twenty characters. - [SceneMixer on voice source: presets, or a recording](https://sentioscope.com/fields/voice-source/scenemixer/): SceneMixer names two routes, a preset voice library or the user's own recordings, and the sentence sits in a compliance checklist, not in product copy. - [sync-3 on voice source: nothing to hand over](https://sentioscope.com/fields/voice-source/sync-3/): sync-3 lists audio and text among its inputs and discusses no voice, so the largest language count in the register comes with no casting mechanism. - [Synthesia on voice source: a catalogue, or a clone](https://sentioscope.com/fields/voice-source/synthesia/): Synthesia publishes a voice catalogue with a language, gender, name and id on every entry, and a cloning route carrying a language list of its own. - [Wan 3.0 on voice source: fifteen seconds, in writing](https://sentioscope.com/fields/voice-source/wan-3/): Wan 3.0 publishes reference audio limits in full: WAV or MP3, one to fifteen seconds a clip, fifteen seconds in total and no more than 15 MB. - [Wan2.2-S2V on voice source: the supplied track](https://sentioscope.com/fields/voice-source/wan-s2v/): The Wan2.2-S2V card takes an audio input with a reference image, so the voice is whatever recording a production supplies, with no catalogue at all. - [D-ID on binding: a voice id, request by request](https://sentioscope.com/fields/per-character-binding/d-id/): D-ID selects the voice as an id on each talk request, with an optional language beside it, and describes nothing that holds the pairing between calls. - [Hedra on binding: whatever travels with the call](https://sentioscope.com/fields/per-character-binding/hedra/): Hedra sends audio or a voice id with each call and publishes nothing about a voice belonging to a character across generations. - [HeyGen on binding: a clone, reused by its own id](https://sentioscope.com/fields/per-character-binding/heygen/): HeyGen enrols a voice clone once, instant or professional, and the clone is then passed as a voice id, so the pairing survives every later request. - [Kling AI on binding: the voice lives on the element](https://sentioscope.com/fields/per-character-binding/kling-ai/): Kling AI binds voices to elements, so a character carries its voice between generations without anything being re-specified on each call. - [LTX Studio on binding: attached to the character](https://sentioscope.com/fields/per-character-binding/ltx-studio/): LTX Studio attaches voices to character Elements, so the voice sits in the same place as the character's appearance and needs no per-shot specification. - [MiniMax on binding: the clip goes in every time](https://sentioscope.com/fields/per-character-binding/minimax/): MiniMax re-establishes a voice from reference audio on each call, capped at fifteen seconds across three clips, with no way published to store it. - [PixVerse on binding: a speaker id, per generation](https://sentioscope.com/fields/per-character-binding/pixverse/): PixVerse passes the chosen speaker id as the lip-sync speaker on each generation request, and publishes nothing about that pairing outlasting the call. - [Runway on binding: a stored voice with its own id](https://sentioscope.com/fields/per-character-binding/runway/): Runway creates a voice asynchronously, brings it to a ready state with a preview, and then addresses it by its own id on later requests. - [SceneMixer on binding: a library, no stated pairing](https://sentioscope.com/fields/per-character-binding/scenemixer/): SceneMixer names a preset library and the user's own recordings as voice routes, and describes nothing about a chosen voice staying with a character. - [Sora 2 on binding: a label inside one prompt](https://sentioscope.com/fields/per-character-binding/sora-2/): Sora 2 asks for speakers to be labelled consistently with turns alternated, which routes a line inside one generation and carries no identity beyond it. - [Synthesia on binding: two objects, no stated pairing](https://sentioscope.com/fields/per-character-binding/synthesia/): Synthesia publishes voice ids and documents avatars, and never states how a particular voice is paired with a particular avatar across videos. - [Wan 3.0 on binding: reference audio, per generation](https://sentioscope.com/fields/per-character-binding/wan-3/): Wan 3.0 supplies reference audio on each generation, within fifteen seconds in total, and publishes nothing about a voice being stored between calls. - [D-ID on lip-sync: a talk, with no mouth described](https://sentioscope.com/fields/lip-sync/d-id/): D-ID's create-a-talk reference documents scripts, audio ceilings and five voice providers, and never says what a generated mouth is following. - [Hedra on lip-sync: it follows the audio provided](https://sentioscope.com/fields/lip-sync/hedra/): Hedra states that the character in the image will lip-sync and move to the audio provided, which names the driver rather than the feature. - [HeyGen on lip-sync: named beside translation](https://sentioscope.com/fields/lip-sync/heygen/): HeyGen names lip-sync alongside voice cloning in its description of video translation, and gives nothing about what drives the mouth. - [Kling AI on lip-sync: a capability, no mechanism](https://sentioscope.com/fields/lip-sync/kling-ai/): Kling AI names lip-sync as a capability of the model and states nothing about what drives it, or what happens with two speakers in frame. - [MiniMax on lip-sync: tied to the speaker on screen](https://sentioscope.com/fields/lip-sync/minimax/): MiniMax ties lip movement to the speaker who is on screen, which addresses the two-speaker case as a behaviour rather than as a control. - [PixVerse on lip-sync: the endpoint's whole purpose](https://sentioscope.com/fields/lip-sync/pixverse/): PixVerse puts lip sync in the name of the endpoint and passes a speaker id as the lip-sync speaker, without saying what drives the movement. - [SceneMixer on lip-sync: a step argued out of existence](https://sentioscope.com/fields/lip-sync/scenemixer/): SceneMixer states there is no separate dubbing step and no lip-sync patch, which answers this column by removing the stage rather than describing it. - [Sora 2 on lip-sync: named where it stops holding](https://sentioscope.com/fields/lip-sync/sora-2/): Sora 2's prompting guide warns that long, complex speeches are unlikely to sync, which names the failure without naming a mechanism. - [sync-3 on lip-sync: matched to the audio it is given](https://sentioscope.com/fields/lip-sync/sync-3/): sync-3 matches lip movement to the audio supplied on the call, which is the clearest driver statement in this column because the audio is the input. - [Synthesia on lip-sync: from the script, framed close](https://sentioscope.com/fields/lip-sync/synthesia/): Synthesia generates lip sync and facial expressions from the spoken content, and states that performance improves with the avatar framed closer to camera. - [Wan 3.0 on lip-sync: audio by default, no mouths](https://sentioscope.com/fields/lip-sync/wan-3/): Wan 3.0 publishes audio defaults, reference caps, pricing and a dialogue syntax, and says nothing anywhere about what a generated mouth follows. - [Wan2.2-S2V on lip-sync: synchronised to the audio](https://sentioscope.com/fields/lip-sync/wan-s2v/): The Wan2.2-S2V card states the result stays synchronised to the audio input, even while a pose sequence is driving the body. - [Runway on lip-sync: voices defined, mouths not](https://sentioscope.com/fields/lip-sync/runway/): Runway's custom voices reference builds voices with samples, previews and ids, and says nothing about a mouth or a performance anywhere on the page. ## Routes to a voice - [Entries that make the sound while they make the shot](https://sentioscope.com/routes/audio-with-the-picture/): Eight entries generate audio in the same pass as the picture, which means alignment arrives for free and the words are the part that resists direction. - [Entries handed a recording and asked to fit a picture](https://sentioscope.com/routes/audio-handed-over/): Three entries take audio as an input and build the picture around it, which puts casting, a session and a delivered file in front of the first frame. - [Entries that synthesise speech before any picture](https://sentioscope.com/routes/speech-from-a-script/): Four entries turn a script into speech as a step of its own, which makes the audio an artefact that can be fetched, approved and reused before rendering. - [Entries whose pages never mention sound at all](https://sentioscope.com/routes/nothing-published/): Two entries reach every field in this register empty, one because its reference is about pictures and one because its reference is about voices. - [Entries where a voice exists before any shot does](https://sentioscope.com/routes/a-stored-voice/): Four entries let a voice be created, named and reused, so a character's voice is wrong once if it is wrong at all rather than on every generation. - [Entries that rebuild the voice on every generation](https://sentioscope.com/routes/a-voice-per-call/): Five entries establish the voice on the request that uses it, by a clip, an id or a label, and publish nothing that holds the pairing afterwards. - [Entries that publish language names, not a total](https://sentioscope.com/routes/languages-named/): Two entries answer the language question with names rather than a number: a claim a production can check, instead of one it has to believe. - [Entries that answer with a number and no names](https://sentioscope.com/routes/languages-counted/): Three entries publish a total for spoken languages, from five to ninety-five or more, and none of them says which languages the total covers. - [Entries claiming languages without naming or counting](https://sentioscope.com/routes/languages-claimed/): Three entries assert more than one language and give neither a list nor a total, which is the weakest grade this register records as an answer. - [Entries that publish a limit as a number](https://sentioscope.com/routes/a-published-figure/): Nine entries put a figure in writing: seconds of reference audio, minutes of enrolment, megabytes per file or the length of a single generation. - [Entries that name what a generated mouth follows](https://sentioscope.com/routes/a-mouth-with-a-driver/): Five entries say what the mouth is following rather than that lip-sync exists, which is the difference between planning coverage and guessing at it. - [Entries that name lip-sync and leave the driver out](https://sentioscope.com/routes/a-mouth-named-only/): Four entries put lip-sync on the page without saying what moves the mouth, which leaves a reader unable to predict which of three failures to expect. - [Entries whose pages never raise the mouth at all](https://sentioscope.com/routes/no-mouth-at-all/): Seven entries reach the lip-sync field empty, and the pattern in the blanks says more about what API references are for than about the products. ## Two entries side by side - [Wan 3.0 and Wan2.2-S2V: two opposite architectures](https://sentioscope.com/against/wan-3-and-wan-s2v/): Two entries from the same model family sit at opposite ends of this register: one makes the sound with the picture, the other is driven by a recording. - [Veo and Hedra: sound out of one, into the other](https://sentioscope.com/against/veo-and-hedra/): One entry generates audio with the video and documents nothing else; the other cannot run without a recording and names what the mouth follows. - [Vidu and sync-3: three tracks made, or one matched](https://sentioscope.com/against/vidu-and-sync-3/): One entry generates speech, effects and music in one pass; the other takes a waveform and moves a mouth to it, and publishes a driver for doing so. - [Synthesia and sync-3: a catalogue against a count](https://sentioscope.com/against/synthesia-and-sync-3/): One entry publishes a language code on every voice; the other publishes ninety-five or more as a total and never mentions a voice at all. - [SceneMixer and Kling AI: names against a count of five](https://sentioscope.com/against/scenemixer-and-kling-ai/): Both generate audio with the picture. One names fifteen dialogue languages and a sixteenth; the other counts five and names none of them. - [Runway and MiniMax: a stored voice, or a clip each time](https://sentioscope.com/against/runway-and-minimax/): One entry builds a voice once, previews it and reuses it by id; the other rebuilds it from fifteen seconds of audio on every generation. - [HeyGen and LTX Studio: a voice object, or a character](https://sentioscope.com/against/heygen-and-ltx-studio/): Both let a voice outlast a call. One enrols a clone with a published threshold in minutes; the other attaches a voice to a character and names no source. - [MiniMax and Wan 3.0: the same fifteen seconds, differently](https://sentioscope.com/against/minimax-and-wan-3/): Two entries publish an identical fifteen-second total for reference audio and arrange everything around it differently, from clip counts to pricing. - [Sora 2 and Luma Ray: a dialogue guide, or silence](https://sentioscope.com/against/sora-2-and-luma-ray/): One entry publishes where a line goes, how to label speakers and how much speech fits a shot; the other has a reference that never mentions sound. - [Synthesia and D-ID: a driver named, or never mentioned](https://sentioscope.com/against/synthesia-and-d-id/): Two talking-head products. One says lip sync comes from the spoken content and adds a framing condition; the other never mentions a mouth. - [HeyGen and Synthesia: a count, or a code per voice](https://sentioscope.com/against/heygen-and-synthesia/): Two avatar platforms answering the language question at different grades: thirty or more counted on translation, against a code published on every voice. - [D-ID and Hedra: a script read out, or a take followed](https://sentioscope.com/against/d-id-and-hedra/): Two avatar products taking opposite inputs: one reads a script of up to forty thousand characters, the other requires a recording and follows its length. - [Kling AI and Synthesia: a capability, or a mechanism](https://sentioscope.com/against/kling-ai-and-synthesia/): Both name lip-sync. One leaves the driver out entirely; the other says it comes from the spoken content and attaches a framing condition. - [SceneMixer and Synthesia: two ways to publish a list](https://sentioscope.com/against/scenemixer-and-synthesia/): Both entries that name languages here do it in a different shape: a per-project setting of fifteen, against a code published on every voice. - [PixVerse and D-ID: one id, two supply chains](https://sentioscope.com/against/pixverse-and-d-id/): Both pass a voice identifier on each request. One draws it from built-in and custom voices; the other from five named outside providers. - [Sora 2 and Wan 3.0: two ways to write a spoken line](https://sentioscope.com/against/sora-2-and-wan-3/): Both put dialogue into the prompt and disagree on how. One wants a labelled block under the prose; the other wants the line written after the word saying. - [Runway and Luma Ray: two empty rows, two reasons](https://sentioscope.com/against/runway-and-luma-ray/): Both entries reach most fields empty. One publishes a thorough voices reference with no shot in it; the other a video reference with no sound in it. - [LTX Studio and Vidu: a bound voice, or none described](https://sentioscope.com/against/ltx-studio-and-vidu/): Both generate audio with the picture. One attaches voices to character Elements; the other describes three tracks and identifies nobody as speaking. - [MiniMax and Sora 2: two speakers in one frame](https://sentioscope.com/against/minimax-and-sora-2/): Two native-audio entries that say something about a shot with two people in it, one as a behaviour and one as an instruction to the writer. - [Hedra and Synthesia: a waveform, or the script itself](https://sentioscope.com/against/hedra-and-synthesia/): Two avatar products naming a driver for the mouth. One follows the audio provided; the other generates movement from the spoken content. ## Section - [Speech controls as a CSV](https://sentioscope.com/datasets/): Recorded speech and voice controls as a comma-separated file, each row carrying the documentation page it came from and the reading date. - [Speech controls across models](https://sentioscope.com/controls/): The comparison table for documented speech controls: where audio comes from, which languages are named, and what each model says about timing. - [Model notes on speech](https://sentioscope.com/models/): A note per model: whether audio comes out with the picture, what can be supplied as a reference, and which languages are named for dialogue. - [The fields behind each model note](https://sentioscope.com/fields/): One page per field in the speech table: what the column records, what wording filled it, and what a filled cell still leaves unsettled. - [Routes to a spoken performance](https://sentioscope.com/routes/): Entries collected by how they answer one column: sound made with the picture, handed over, synthesised from a script, or never mentioned. - [Entries read side by side](https://sentioscope.com/against/): Pairs of entries put on one page because at least one column lands them at opposite grades of answer, with the gap named on every field. - [Questions about generated speech](https://sentioscope.com/questions/): Questions about dialogue in AI video, answered from vendor documentation rather than from listening tests, with the wording quoted and dated. - [Speech terms in these notes](https://sentioscope.com/terms/): Definitions for the speech vocabulary these notes use, from native audio and reference clips to voice identifiers, language codes and lip-sync. - [Explainers on generated speech](https://sentioscope.com/learn/): Background notes on dialogue in generative video: how lip-sync is produced, where timing breaks, and what a voice reference actually promises.