D-ID: a script read aloud, or a file handed over
The create-a-talk reference gives two ways in: a text script capped at forty thousand characters, or an audio url. Supplied audio runs to five minutes for a clip and ten for a talk, so the ceiling is written down rather than discovered. As of 2026-09-22.
| What the documentation settles | What it leaves to a take |
|---|---|
| A script is either text of up to forty thousand characters, or an audio url | Which of the two produces the steadier read of the same line |
| Supplied audio runs to five minutes for a clip, ten for a talk | Whether a face holds attention across ten unbroken minutes |
| Five speech providers are named for the voice | Which of the five a given voice id came from, before it is called |
Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.
1Forty thousand characters is a ceiling nobody reaches
A forty-thousand-character script is roughly five hours of speech, which no single talking-head render is going to hold. Publishing the number anyway is useful in a different way: it says the text route is not a demo affordance with a hidden sentence limit, and a production can paste a whole scene into it without splitting the file first.
The audio route carries the ceiling that actually binds. Five minutes for a clip and ten for a talk are numbers a schedule can be built from, and they are the reason a long piece here is assembled from several calls rather than requested in one.
2Two routes that fail in opposite directions
Handing over a recording fixes the performance and leaves the picture to follow. Handing over text fixes the words and leaves the performance to a synthesiser that a director never auditioned. Neither is better in the abstract; they put the risk in different places, and only one of them can be approved before the render is paid for.
Because both routes land on the same endpoint, the choice is easy to make by accident. A team prototyping with text and delivering with recordings is testing something other than what it will ship, and nothing in the returned file says which route produced it.
3Others that make speech a step of its own
Three more entries here treat speech as a stage before the picture rather than as something the picture arrives with. What differs between them is how much of that stage is published.
- HeyGen — from a speech endpoint, then rendered.
- PixVerse — supplied, or read from text.
- Synthesia — from the script, or uploaded.
- Audio sourceA script is either text of up to forty thousand characters or an audio url, with audio limited to five minutes for clips and ten for talksa script read aloud or a file
- Voice sourceFive speech providers are named for the voice: Microsoft, ElevenLabs, Amazon, Google and Azure OpenAIoutside catalogues named on the page
4Sources
Read from the create a talk reference at docs.d-id.com on 2026-09-22. The same column across every entry is on audio source; everything this vendor publishes about speech is on D-ID. What counts as documented is on how read.