D-ID: the voice is bought in, and the sellers are named
Five speech providers are named on the reference: Microsoft, ElevenLabs, Amazon, Google and Azure OpenAI. The voice is a voice id chosen from one of those catalogues, which puts the casting decision one level outside the product. As of 2026-09-22.
| What the documentation settles | What it leaves to a take |
|---|---|
| Five speech providers are named for the voice | Which of the five a production should prefer for a dramatic read |
| The voice is a voice id selected from the list of available voices | How large that list is, and how it is browsed |
| A script is either text of up to forty thousand characters or an audio url | Whether a supplied recording bypasses the providers entirely |
Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.
1Naming the suppliers is more candid than most pages here
Several entries describe a voice as though the platform had built it. This one says plainly that the voices come from five outside services, which lets a reader take the question somewhere it can actually be answered: those services publish their own catalogues, licences and language ranges.
It also makes the pricing conversation legible. A voice that belongs to another vendor usually carries that vendor's terms, and a production planning a season of dialogue is committing to a supply chain rather than to one product.
2A supply chain is a set of dependencies nobody logged
Five providers means five deprecation schedules, five sets of language codes, and five ways a voice can quietly stop existing. None of that reaches a production through this reference, because the reference describes how to pass a voice id and not where the id came from.
The practical habit is to record the provider alongside the voice id in a production's own notes. The returned file does not say which service spoke, and a year later that is the one thing needed to reproduce a character.
3The others that hand a production a catalogue
Four more entries here answer this column with something to choose from rather than something to supply. They differ on whether the catalogue is the vendor's own, and on whether a recording is accepted alongside it.
- Hedra — a track, or a voice id from the endpoint.
- PixVerse — a sample, or a built-in voice.
- SceneMixer — a preset library, or the user's own recordings.
- Synthesia — a catalogue voice, or a cloned one.
- Voice sourceFive speech providers are named for the voice: Microsoft, ElevenLabs, Amazon, Google and Azure OpenAIoutside catalogues named on the page
- Per-character bindingThe voice is a voice id selected from the list of available voices, with an optional language field beside ita voice id on every request
4Sources
Read from the create a talk reference at docs.d-id.com on 2026-09-22. The same column across every entry is on voice source; everything this vendor publishes about speech is on D-ID. What counts as documented is on how read.