Sentioscope

Speech and voice controls, as each vendor documents them

What Sora 2 documents about speech

Sora 2 (developers.openai.com) is described as generating videos with synced audio, and its prompting guide is the one place in this register that tells a writer where a line of dialogue goes and how to label two people talking. As of 2026-09-22.

Where a spoken line goes in a request, and what sizes itThe guidance splits a request in two: prose describing what is seen, then a separate block holding the words. Speakers are labelled inside that block so lines land on the right face, and the number of exchanges is governed by how long the clip runs.Prose firstWhat the shotlooks like, inordinarysentences.ThenDialogue blockThe spokenlines, keptapart from thedescription.ThenSpeaker labelsNamedconsistently,with turnsalternatingbetween them.Sized byClip lengthOne or twoshort exchangesin fourseconds; a fewmore in eight.
Fig. 1 Every step here is a published instruction about writing, which is rarer in this register than a claim about sound.
Controls documented, model by modelFilled where the model documents that control, hollow where nothing is published about it.Controls documented, model by modelWhat the vendor documentsWhat the vendordocumentsAudio sourceAudio source — What the vendor documents: Generated with the picture; video and audio are both listed as outputLanguagesLanguages — What the vendor documents: Not documentedVoice sourceVoice source — What the vendor documents: Not documentedPer-character bindingPer-character binding — What the vendor documents: Not documented; speakers are labelled inside the promptLip-syncLip-sync — What the vendor documents: Not named; long speeches are said to be unlikely to sync
Fig. 2 Filled where the model documents that control, hollow where nothing is published about it.
The five fields this register keeps, for Sora 2. Read from the vendor's model page and prompting guide on 2026-09-22.
FieldWhat the vendor documents
Audio sourceGenerated with the picture; video and audio are both listed as output
LanguagesNot documented
Voice sourceNot documented
Per-character bindingNot documented; speakers are labelled inside the prompt
Lip-syncNot named; long speeches are said to be unlikely to sync

Inclusion rule. Fields are the same five for every model here; a field the vendor does not document is recorded as not documented rather than inferred from generated output. Order. Fixed field order, identical on every model page.

1Dialogue has a place in the prompt, and that is unusual

Most entries here describe audio as an output and stop there. This one tells a writer where the words go: spoken lines sit in a dialogue block underneath the prose description, so the model can tell a description of a room from a line said inside it. Exchanges are asked to stay within a handful of sentences.

For a shot list that is a format instruction rather than a quality claim. It means dialogue in a production's own documents can be written in the shape the model expects, and the conversion step that usually sits between a script and a prompt mostly disappears.

2Arithmetic a writer can plan against

The guide puts one or two short exchanges in a four-second shot and a few more in an eight-second one, and warns that long, complex speeches are unlikely to sync. Read beside the 16 and 20 second generation lengths, and the extension that adds up to 20 seconds at a time to a total of 120, that is enough to size a scene before anybody spends a call on it.

It is also as close as the page comes to a statement about mouths: not a mechanism, but an admission of where alignment stops holding. Recording lip-sync as documented on that basis would be over-reading it, so the field stays marked unstated.

3Whose voice it is stays unaddressed

Nothing on these pages says where a voice comes from, whether one can be supplied, or whether a character keeps one between calls. Speaker labels route lines to the right face inside a single generation; a label is an instruction within one call, not an identity that survives the next.

For a serial that is the gap that matters. A production can write a two-hander that plays correctly in one shot and still have no published way to make the same character sound the same in the following episode.

4This entry, one column at a time

Each of these stays inside a single field: what Sora 2 puts there, what the wording settles, and what it leaves for a take to answer.

5Read against another entry

Each of these puts Sora 2 beside one other entry on a column where the two land at opposite grades of answer.

6Where this model sits against the rest

Audio source, across every model hereOne field at a time, every entry side by side, with Sora 2 marked. The full wording for every cell is in the rows below. This entry was read 2026-09-22.Audio source, across every model hereAudio sourceD-IDD-ID — Audio source: A text script read aloud, or an audio url suppliedHedraHedra — Audio source: Supplied; audio is a required inputHeyGenHeyGen — Audio source: A text to speech endpoint, apart from the pictureKling AIKling AI — Audio source: Generated with the picture, VIDEO 3.0LTX StudioLTX Studio — Audio source: Generated jointly with the picture, plus audio-to-videoLuma RayLuma Ray — Audio source: Not documented; the video reference does not mention soundMiniMaxMiniMax — Audio source: Native speech, generated with the picturePixVersePixVerse — Audio source: Supplied to a speech endpoint, or read from text thereRunwayRunway — Audio source: Not documented where voices are definedSceneMixerSceneMixer — Audio source: Generated with the picture, in the language set for the projectSora 2Sora 2 — Audio source: Generated with the picture; video and audio are both listed as outputsync-3sync-3 — Audio source: Supplied, or read from text on the callSynthesiaSynthesia — Audio source: Synthesised from the script, or an uploaded recordingVeoVeo — Audio source: Generated with the video rather than added afterwardsViduVidu — Audio source: Q3 generates speech, effects and music natively, on by defaultWan 3.0Wan 3.0 — Audio source: On by default; false returns a file with no audio trackWan2.2-S2VWan2.2-S2V — Audio source: Supplied; the card describes the model as audio-driven
Fig. 3 One field at a time, every entry side by side, with Sora 2 marked. The full wording for every cell is in the rows below. This entry was read 2026-09-22.
Every entry in this register on the same field, with Sora 2 marked. Each cell carries the reading date of its own entry; this one was read 2026-09-22.
ModelAudio source
D-IDA text script read aloud, or an audio url supplied
HedraSupplied; audio is a required input
HeyGenA text to speech endpoint, apart from the picture
Kling AIGenerated with the picture, VIDEO 3.0
LTX StudioGenerated jointly with the picture, plus audio-to-video
Luma RayNot documented; the video reference does not mention sound
MiniMaxNative speech, generated with the picture
PixVerseSupplied to a speech endpoint, or read from text there
RunwayNot documented where voices are defined
SceneMixerGenerated with the picture, in the language set for the project
Sora 2Generated with the picture; video and audio are both listed as output
sync-3Supplied, or read from text on the call
SynthesiaSynthesised from the script, or an uploaded recording
VeoGenerated with the video rather than added afterwards
ViduQ3 generates speech, effects and music natively, on by default
Wan 3.0On by default; false returns a file with no audio track
Wan2.2-S2VSupplied; the card describes the model as audio-driven

Inclusion rule. Models whose vendor documents something about generated speech. A model with nothing published on the point is still listed, with the cell marked hollow. Order. Alphabetical by model name.

Languages, across every model hereOne field at a time, every entry side by side, with Sora 2 marked. The full wording for every cell is in the rows below. This entry was read 2026-09-22.Languages, across every model hereLanguagesD-IDD-ID — Languages: A language field beside the voice; no list on the referenceHedraHedra — Languages: Full multi-language support claimed, none namedHeyGenHeyGen — Languages: Thirty or more, counted on the translation endpointKling AIKling AI — Languages: Five, not namedLTX StudioLTX Studio — Languages: Not documentedLuma RayLuma Ray — Languages: Not documentedMiniMaxMiniMax — Languages: Not documented by the vendorPixVersePixVerse — Languages: Multiple, with speech, singing and advertisements named as typesRunwayRunway — Languages: Not documentedSceneMixerSceneMixer — Languages: 15, named, with Cantonese a 16th for dialogue onlySora 2Sora 2 — Languages: Not documentedsync-3sync-3 — Languages: Ninety-five or more, counted rather than namedSynthesiaSynthesia — Languages: A catalogue, each voice carrying a language name and codeVeoVeo — Languages: Not documentedViduVidu — Languages: Not documentedWan 3.0Wan 3.0 — Languages: Not documentedWan2.2-S2VWan2.2-S2V — Languages: Not documented
Fig. 4 One field at a time, every entry side by side, with Sora 2 marked. The full wording for every cell is in the rows below. This entry was read 2026-09-22.
Every entry in this register on the same field, with Sora 2 marked. Each cell carries the reading date of its own entry; this one was read 2026-09-22.
ModelLanguages
D-IDA language field beside the voice; no list on the reference
HedraFull multi-language support claimed, none named
HeyGenThirty or more, counted on the translation endpoint
Kling AIFive, not named
LTX StudioNot documented
Luma RayNot documented
MiniMaxNot documented by the vendor
PixVerseMultiple, with speech, singing and advertisements named as types
RunwayNot documented
SceneMixer15, named, with Cantonese a 16th for dialogue only
Sora 2Not documented
sync-3Ninety-five or more, counted rather than named
SynthesiaA catalogue, each voice carrying a language name and code
VeoNot documented
ViduNot documented
Wan 3.0Not documented
Wan2.2-S2VNot documented

Inclusion rule. Models whose vendor documents something about generated speech. A model with nothing published on the point is still listed, with the cell marked hollow. Order. Alphabetical by model name.

Voice source, across every model hereOne field at a time, every entry side by side, with Sora 2 marked. The full wording for every cell is in the rows below. This entry was read 2026-09-22.Voice source, across every model hereVoice sourceD-IDD-ID — Voice source: A voice id from one of five named speech providersHedraHedra — Voice source: An uploaded track, or a voice id from the voices endpointHeyGenHeyGen — Voice source: A clone, instant from one recording or professional from 20 minutes or moreKling AIKling AI — Voice source: Not documented by the vendorLTX StudioLTX Studio — Voice source: Not documentedLuma RayLuma Ray — Voice source: Not documentedMiniMaxMiniMax — Voice source: Reference audio, 15 s total across 3 clipsPixVersePixVerse — Voice source: Built-in voices, or a custom voice from a supplied sampleRunwayRunway — Voice source: An audio sample of 10 seconds to 5 minutes, or a description in wordsSceneMixerSceneMixer — Voice source: Preset voice library, or the user's own recordingsSora 2Sora 2 — Voice source: Not documentedsync-3sync-3 — Voice source: Not documented on the model pageSynthesiaSynthesia — Voice source: The catalogue, or a cloned voice with a language list of its ownVeoVeo — Voice source: Not documentedViduVidu — Voice source: Not documentedWan 3.0Wan 3.0 — Voice source: Reference audio, 15 seconds in total, WAV or MP3 up to 15 MBWan2.2-S2VWan2.2-S2V — Voice source: Whatever track is handed over; no catalogue
Fig. 5 One field at a time, every entry side by side, with Sora 2 marked. The full wording for every cell is in the rows below. This entry was read 2026-09-22.
Every entry in this register on the same field, with Sora 2 marked. Each cell carries the reading date of its own entry; this one was read 2026-09-22.
ModelVoice source
D-IDA voice id from one of five named speech providers
HedraAn uploaded track, or a voice id from the voices endpoint
HeyGenA clone, instant from one recording or professional from 20 minutes or more
Kling AINot documented by the vendor
LTX StudioNot documented
Luma RayNot documented
MiniMaxReference audio, 15 s total across 3 clips
PixVerseBuilt-in voices, or a custom voice from a supplied sample
RunwayAn audio sample of 10 seconds to 5 minutes, or a description in words
SceneMixerPreset voice library, or the user's own recordings
Sora 2Not documented
sync-3Not documented on the model page
SynthesiaThe catalogue, or a cloned voice with a language list of its own
VeoNot documented
ViduNot documented
Wan 3.0Reference audio, 15 seconds in total, WAV or MP3 up to 15 MB
Wan2.2-S2VWhatever track is handed over; no catalogue

Inclusion rule. Models whose vendor documents something about generated speech. A model with nothing published on the point is still listed, with the cell marked hollow. Order. Alphabetical by model name.

7Sources

Taken from the model page and prompting guide at developers.openai.com on 2026-09-22. Every entry against the same five fields sits on the speech controls page; the rule behind each field is on how read. The entries either side of this one: sync-3, Synthesia.

  • Audio source
    Described as a media generation model generating videos with synced audio, with video and audio both listed as outputgenerated with the pictureOpenAI, Sora 2 model page / recorded 2026-09-22
  • Shot length a line has to fit
    Generations run 16 or 20 seconds, and an extension adds up to 20 seconds at a time to a total of 120OpenAI, video generation guide / recorded 2026-09-22
  • How a spoken line is written into a prompt
    Dialogue belongs in a dialogue block below the prose description, with exchanges limited to a handful of sentencesOpenAI, Sora 2 prompting guide / recorded 2026-09-22
  • Two characters in frame
    For multi-character scenes the guide asks for speakers labelled consistently and turns alternated, so each line lands on the right characteraddressed rather than skippedOpenAI, Sora 2 prompting guide / recorded 2026-09-22
  • Dialogue timing against clip length
    A four-second shot is said to accommodate one or two short exchanges and an eight-second clip a few more, with long speeches unlikely to syncOpenAI, Sora 2 prompting guide / recorded 2026-09-22
  • Languages
    Not documented by the vendor (as of 2026-09-22)neither a list nor a countOpenAI, video generation guide / recorded 2026-09-22