What D-ID documents about speech
D-ID (docs.d-id.com) takes either a text script of up to forty thousand characters or an audio url, names five speech providers a voice can come from, and limits supplied audio to five minutes for a clip and ten for a talk. As of 2026-09-22.
| Field | What the vendor documents |
|---|---|
| Audio source | A text script read aloud, or an audio url supplied |
| Languages | A language field beside the voice; no list on the reference |
| Voice source | A voice id from one of five named speech providers |
| Per-character binding | A voice id on every request |
| Lip-sync | Not documented on the reference |
Inclusion rule. Fields are the same five for every model here; a field the vendor does not document is recorded as not documented rather than inferred from generated output. Order. Fixed field order, identical on every model page.
1The voice is bought in, and the reference says from whom
Microsoft, ElevenLabs, Amazon, Google and Azure OpenAI are named as the providers a voice may come from. Naming them is more informative than a house catalogue would be, because a production can look the voice up at its source and know what it is getting before committing a character to it.
It also makes a dependency visible. A voice belongs to a provider, and a provider can revise its catalogue, so a series that pinned a lead to one id has a supply chain behind that character rather than a setting.
2Text or a recording, with the ceilings written down
A script can run to forty thousand characters, ten thousand of them outside markup, or a recording arrives by url. Supplied audio is limited to five minutes for a clip and ten for a talk. Those are generous next to the single-figure ceilings elsewhere in this register, and they describe a different product: a presenter delivering a script rather than a shot inside a scene.
For drama the useful reading is the shape of the limit. Long-form speech is expected here, so the timing problem that dominates generated shots barely arises, and framing takes its place as the thing to get right.
3What the reference leaves to some other page
A language field sits beside the voice, with a formal name as its example value, but the reference read for this entry carries neither a list nor a count of languages, and says nothing about what drives the mouth. Both may be documented elsewhere on the platform; this register records the page it read and dates it.
Recording the gap instead of hunting for a friendlier page is deliberate. The question a production asks is whether the thing it needs is where it would look, and the answer here is that the voice is and the language reach is not.
4This entry, one column at a time
Each of these stays inside a single field: what D-ID puts there, what the wording settles, and what it leaves for a take to answer.
- Audio source — from a script, or a supplied url.
- Voice source — a voice id, or a recording by url.
- Per-character binding — on the request.
- Languages — a language field, no list.
- Lip-sync — not documented by the vendor.
5Read against another entry
Each of these puts D-ID beside one other entry on a column where the two land at opposite grades of answer.
- Synthesia and D-ID — on lip-sync.
- D-ID and Hedra — on audio source.
- PixVerse and D-ID — on per-character binding.
6Where this model sits against the rest
| Model | Audio source |
|---|---|
| D-ID | A text script read aloud, or an audio url supplied |
| Hedra | Supplied; audio is a required input |
| HeyGen | A text to speech endpoint, apart from the picture |
| Kling AI | Generated with the picture, VIDEO 3.0 |
| LTX Studio | Generated jointly with the picture, plus audio-to-video |
| Luma Ray | Not documented; the video reference does not mention sound |
| MiniMax | Native speech, generated with the picture |
| PixVerse | Supplied to a speech endpoint, or read from text there |
| Runway | Not documented where voices are defined |
| SceneMixer | Generated with the picture, in the language set for the project |
| Sora 2 | Generated with the picture; video and audio are both listed as output |
| sync-3 | Supplied, or read from text on the call |
| Synthesia | Synthesised from the script, or an uploaded recording |
| Veo | Generated with the video rather than added afterwards |
| Vidu | Q3 generates speech, effects and music natively, on by default |
| Wan 3.0 | On by default; false returns a file with no audio track |
| Wan2.2-S2V | Supplied; the card describes the model as audio-driven |
Inclusion rule. Models whose vendor documents something about generated speech. A model with nothing published on the point is still listed, with the cell marked hollow. Order. Alphabetical by model name.
| Model | Languages |
|---|---|
| D-ID | A language field beside the voice; no list on the reference |
| Hedra | Full multi-language support claimed, none named |
| HeyGen | Thirty or more, counted on the translation endpoint |
| Kling AI | Five, not named |
| LTX Studio | Not documented |
| Luma Ray | Not documented |
| MiniMax | Not documented by the vendor |
| PixVerse | Multiple, with speech, singing and advertisements named as types |
| Runway | Not documented |
| SceneMixer | 15, named, with Cantonese a 16th for dialogue only |
| Sora 2 | Not documented |
| sync-3 | Ninety-five or more, counted rather than named |
| Synthesia | A catalogue, each voice carrying a language name and code |
| Veo | Not documented |
| Vidu | Not documented |
| Wan 3.0 | Not documented |
| Wan2.2-S2V | Not documented |
Inclusion rule. Models whose vendor documents something about generated speech. A model with nothing published on the point is still listed, with the cell marked hollow. Order. Alphabetical by model name.
| Model | Voice source |
|---|---|
| D-ID | A voice id from one of five named speech providers |
| Hedra | An uploaded track, or a voice id from the voices endpoint |
| HeyGen | A clone, instant from one recording or professional from 20 minutes or more |
| Kling AI | Not documented by the vendor |
| LTX Studio | Not documented |
| Luma Ray | Not documented |
| MiniMax | Reference audio, 15 s total across 3 clips |
| PixVerse | Built-in voices, or a custom voice from a supplied sample |
| Runway | An audio sample of 10 seconds to 5 minutes, or a description in words |
| SceneMixer | Preset voice library, or the user's own recordings |
| Sora 2 | Not documented |
| sync-3 | Not documented on the model page |
| Synthesia | The catalogue, or a cloned voice with a language list of its own |
| Veo | Not documented |
| Vidu | Not documented |
| Wan 3.0 | Reference audio, 15 seconds in total, WAV or MP3 up to 15 MB |
| Wan2.2-S2V | Whatever track is handed over; no catalogue |
Inclusion rule. Models whose vendor documents something about generated speech. A model with nothing published on the point is still listed, with the cell marked hollow. Order. Alphabetical by model name.
7Sources
Taken from the create a talk reference at docs.d-id.com on 2026-09-22. Every entry against the same five fields sits on the speech controls page; the rule behind each field is on how read. The entries either side of this one: Hedra, HeyGen.
- Audio sourceA script is either text of up to forty thousand characters or an audio url, with audio limited to five minutes for clips and ten for talksa script read aloud or a file
- Voice sourceFive speech providers are named for the voice: Microsoft, ElevenLabs, Amazon, Google and Azure OpenAIoutside catalogues named on the page
- Per-character bindingThe voice is a voice id selected from the list of available voices, with an optional language field beside ita voice id on every request