Runway: a sample, or a sentence describing a voice
Two routes sit side by side. A sample of ten seconds to five minutes, at most ten megabytes, or a description in words with a published minimum length rather than a maximum. The second route exists nowhere else in this register. As of 2026-09-22.
| What the documentation settles | What it leaves to a take |
|---|---|
| A custom voice can be built from a sample of 10 seconds to 5 minutes, at most 10 MB | How much of the read a five-minute sample carries that ten seconds does not |
| A voice can instead be described in words, with at least twenty characters | How specific a description has to be to land where it was aimed |
| A voice is created asynchronously, reaches a ready state with a preview, then is used by id | Whether a described voice is reproducible from the same sentence |
Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.
1A minimum length on a description is a revealing constraint
Most published limits are ceilings. This one is a floor: the description has to run at least twenty characters. Read as engineering it stops empty strings; read as a hint it says the model wants enough adjectives to work with, and a three-word description will produce something generic.
The described route also steps around the consent question a sample carries, because there is no identifiable person behind a sentence about a voice. For a minor character nobody was going to record, that is the cheapest casting in this register.
2The sample window is the widest here, which changes the choice
Ten seconds to five minutes spans two different kinds of clip. At the short end this behaves like every other entry in the column, carrying timbre and little else. At five minutes there is room for range, and the choice of what to include becomes a directing decision rather than a technical one.
Nothing on the page says the longer clip produces a better voice, and it is worth not assuming so. A wide window means a production can experiment; it does not mean more audio is always better than a cleaner ten seconds.
3The other entry where a voice is built once and reused
One more entry constructs a voice before any shot and then addresses it by name. It publishes a threshold in minutes of studio time, where this one publishes a window and an alternative that needs no recording at all.
- HeyGen — a recording to clone from.
- Voice sourceA custom voice can be built from an audio sample between 10 seconds and 5 minutes long and at most 10 MBa published sample window
- Voice sourceA voice can instead be asked for in words, with the description required to run at least 20 charactersa route that needs no recording
- Per-character bindingA voice is created asynchronously, reaches a ready state with a preview, and is then referred to by its own ida stored voice object
4Sources
Read from the custom voices reference at docs.dev.runwayml.com on 2026-09-22. The same column across every entry is on voice source; everything this vendor publishes about speech is on Runway. What counts as documented is on how read.