D-ID: a talk gets made, and the mouth goes unmentioned
The endpoint is called create a talk, and the reference never mentions lip movement. Scripts, audio ceilings and five voice providers are all published; what happens to the face while the words play is not. As of 2026-09-22.
| What the documentation settles | What it leaves to a take |
|---|---|
| A script is either text of up to forty thousand characters or an audio url | Whether the mouth follows the audio or the text it came from |
| Five speech providers are named for the voice | Whether alignment differs between provider voices |
| Nothing on the reference addresses lip movement | How the face behaves across a ten-minute talk |
Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.
1A silence that is harder to excuse than most here
Pages written about picture generation can be forgiven for skipping mouths. This one is written about producing a talking person, and the word for the thing the product exists to do does not appear on it. That is the most striking blank in this column.
The likely explanation is that the reference is an API surface rather than a product description: it lists fields a caller sets, and there is no field for a convincing mouth. That explains the omission without making it less consequential for a reader choosing between tools.
2The audio route and the text route may not behave alike
Where a recording is supplied, alignment has a waveform to follow. Where a script is supplied, speech is synthesised first and then presumably followed, which is an extra step and an extra place to lose a frame. Nothing says whether the two produce the same quality.
For a production that difference is testable and worth testing, because the two routes are otherwise interchangeable on this endpoint and a team will drift between them without noticing.
3Others whose pages never raise the subject
Six more entries reach this column with nothing in it. Most of those documents are about generating pictures; this one is about generating a person who talks.
- LTX Studio — not documented by the vendor.
- Luma Ray — not documented by the vendor.
- Runway — not documented by the vendor.
- Veo — not documented by the vendor.
- Vidu — not documented by the vendor.
- Wan 3.0 — not documented by the vendor.
- Audio sourceA script is either text of up to forty thousand characters or an audio url, with audio limited to five minutes for clips and ten for talksa script read aloud or a file
- Voice sourceFive speech providers are named for the voice: Microsoft, ElevenLabs, Amazon, Google and Azure OpenAIoutside catalogues named on the page
- Per-character bindingThe voice is a voice id selected from the list of available voices, with an optional language field beside ita voice id on every request
4Sources
Read from the create a talk reference at docs.d-id.com on 2026-09-22. The same column across every entry is on lip-sync; everything this vendor publishes about speech is on D-ID. What counts as documented is on how read.