What Runway documents about speech
Runway (docs.dev.runwayml.com) documents a custom voice built either from an audio sample of ten seconds to five minutes or from a written description of at least twenty characters, then referred to afterwards by its own id. As of 2026-09-22.
| Field | What the vendor documents |
|---|---|
| Audio source | Not documented where voices are defined |
| Languages | Not documented |
| Voice source | An audio sample of 10 seconds to 5 minutes, or a description in words |
| Per-character binding | A stored voice, referred to by its own id |
| Lip-sync | Not documented |
Inclusion rule. Fields are the same five for every model here; a field the vendor does not document is recorded as not documented rather than inferred from generated output. Order. Fixed field order, identical on every model page.
1A voice can be asked for in words, which nothing else here offers
Two routes sit side by side. One is the familiar sample: ten seconds to five minutes of audio, ten megabytes at most, handed over as a link, an upload or inline data. The other is a text description, and the constraint published on it is a minimum length rather than a maximum, which reads as the platform asking for enough adjectives to work with.
The described route is the one that solves casting a minor character nobody has recorded. It also steps around the consent question a sample carries, because there is no identifiable person behind a sentence about a voice.
2A voice that exists before any shot does
Creation is asynchronous. A voice moves through processing into a ready state, arrives with a preview to listen to, and is then addressed by id. That sequence is what separates a voice from a parameter: it can be auditioned, approved and reused, and the approval happens once instead of inside every generation.
This register records that as a binding with a caveat. An id is a durable handle, but nothing on the page ties one to a character object, so keeping a cast straight remains a production's own bookkeeping.
3What the page settles, and what it does not
Where the sound in a generated shot originates, which languages are spoken, and what drives a mouth are all absent, because this page defines voices rather than rendering them. A model option with multilingual in its name appears; a model name is not a language list, and nothing here promotes one into the other.
What it does settle is the part most vendors leave vague: the exact shape of what a production hands over, written in seconds and megabytes.
4This entry, one column at a time
Each of these stays inside a single field: what Runway puts there, what the wording settles, and what it leaves for a take to answer.
- Voice source — a 10-second to 5-minute sample, or a sentence.
- Per-character binding — on a stored voice.
- Languages — not documented by the vendor.
- Lip-sync — not documented by the vendor.
5Read against another entry
Each of these puts Runway beside one other entry on a column where the two land at opposite grades of answer.
- Runway and MiniMax — on voice source.
- Runway and Luma Ray — on voice source.
6Where this model sits against the rest
| Model | Audio source |
|---|---|
| D-ID | A text script read aloud, or an audio url supplied |
| Hedra | Supplied; audio is a required input |
| HeyGen | A text to speech endpoint, apart from the picture |
| Kling AI | Generated with the picture, VIDEO 3.0 |
| LTX Studio | Generated jointly with the picture, plus audio-to-video |
| Luma Ray | Not documented; the video reference does not mention sound |
| MiniMax | Native speech, generated with the picture |
| PixVerse | Supplied to a speech endpoint, or read from text there |
| Runway | Not documented where voices are defined |
| SceneMixer | Generated with the picture, in the language set for the project |
| Sora 2 | Generated with the picture; video and audio are both listed as output |
| sync-3 | Supplied, or read from text on the call |
| Synthesia | Synthesised from the script, or an uploaded recording |
| Veo | Generated with the video rather than added afterwards |
| Vidu | Q3 generates speech, effects and music natively, on by default |
| Wan 3.0 | On by default; false returns a file with no audio track |
| Wan2.2-S2V | Supplied; the card describes the model as audio-driven |
Inclusion rule. Models whose vendor documents something about generated speech. A model with nothing published on the point is still listed, with the cell marked hollow. Order. Alphabetical by model name.
| Model | Languages |
|---|---|
| D-ID | A language field beside the voice; no list on the reference |
| Hedra | Full multi-language support claimed, none named |
| HeyGen | Thirty or more, counted on the translation endpoint |
| Kling AI | Five, not named |
| LTX Studio | Not documented |
| Luma Ray | Not documented |
| MiniMax | Not documented by the vendor |
| PixVerse | Multiple, with speech, singing and advertisements named as types |
| Runway | Not documented |
| SceneMixer | 15, named, with Cantonese a 16th for dialogue only |
| Sora 2 | Not documented |
| sync-3 | Ninety-five or more, counted rather than named |
| Synthesia | A catalogue, each voice carrying a language name and code |
| Veo | Not documented |
| Vidu | Not documented |
| Wan 3.0 | Not documented |
| Wan2.2-S2V | Not documented |
Inclusion rule. Models whose vendor documents something about generated speech. A model with nothing published on the point is still listed, with the cell marked hollow. Order. Alphabetical by model name.
| Model | Voice source |
|---|---|
| D-ID | A voice id from one of five named speech providers |
| Hedra | An uploaded track, or a voice id from the voices endpoint |
| HeyGen | A clone, instant from one recording or professional from 20 minutes or more |
| Kling AI | Not documented by the vendor |
| LTX Studio | Not documented |
| Luma Ray | Not documented |
| MiniMax | Reference audio, 15 s total across 3 clips |
| PixVerse | Built-in voices, or a custom voice from a supplied sample |
| Runway | An audio sample of 10 seconds to 5 minutes, or a description in words |
| SceneMixer | Preset voice library, or the user's own recordings |
| Sora 2 | Not documented |
| sync-3 | Not documented on the model page |
| Synthesia | The catalogue, or a cloned voice with a language list of its own |
| Veo | Not documented |
| Vidu | Not documented |
| Wan 3.0 | Reference audio, 15 seconds in total, WAV or MP3 up to 15 MB |
| Wan2.2-S2V | Whatever track is handed over; no catalogue |
Inclusion rule. Models whose vendor documents something about generated speech. A model with nothing published on the point is still listed, with the cell marked hollow. Order. Alphabetical by model name.
7Sources
Taken from the custom voices reference at docs.dev.runwayml.com on 2026-09-22. Every entry against the same five fields sits on the speech controls page; the rule behind each field is on how read. The entries either side of this one: SceneMixer, Sora 2.
- Voice sourceA custom voice can be built from an audio sample between 10 seconds and 5 minutes long and at most 10 MBa published sample window
- Voice sourceA voice can instead be asked for in words, with the description required to run at least 20 charactersa route that needs no recording
- Per-character bindingA voice is created asynchronously, reaches a ready state with a preview, and is then referred to by its own ida stored voice object
- Lip-syncNot documented by the vendor (as of 2026-09-22)not addressed where voices are defined