D-ID: a voice id on every request, and nothing above it
The voice is a voice id chosen from the list of available voices, passed on the request that needs it. That makes the pairing between a character and a voice a property of a production's own records rather than of the platform. As of 2026-09-22.
| What the documentation settles | What it leaves to a take |
|---|---|
| The voice is a voice id selected from the list of available voices | Where the pairing of a character to that id is kept |
| An optional language field sits beside the chosen voice | Whether the same character can switch language and keep its voice |
| Five speech providers are named for the voice | Whether an id stays valid if a provider reorganises its catalogue |
Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.
1A stable id is most of what persistence needs
This arrangement is weaker than a stored character object and much stronger than a reference clip. An id is stable, short, and easy to keep in a shot list, so a production that writes it down gets consistency for free and does not have to store audio anywhere.
What it does not give is a guarantee. The platform will not object if the wrong id is passed for a character, and nothing in the returned file records which one was used, so the only defence against a season drifting is a discipline outside the tool.
2Catalogue ids belong to five other companies
Because the voices come from named outside providers, the id is ultimately theirs. A voice deprecated upstream takes a character with it, and the talk reference is not the document that would announce that.
For a series that argues for a short cast list and for recording the provider beside the id. Two fields in a spreadsheet are a cheap insurance policy against a dependency nobody else is tracking.
3The others that re-establish the voice on each call
Four more entries put the voice on the request rather than on the character. Between them they use an id, a reference clip and a label inside the prompt, and the three fail in different ways.
- Per-character bindingThe voice is a voice id selected from the list of available voices, with an optional language field beside ita voice id on every request
- Voice sourceFive speech providers are named for the voice: Microsoft, ElevenLabs, Amazon, Google and Azure OpenAIoutside catalogues named on the page
4Sources
Read from the create a talk reference at docs.d-id.com on 2026-09-22. The same column across every entry is on per-character binding; everything this vendor publishes about speech is on D-ID. What counts as documented is on how read.