D-ID and Hedra: a script read out, or a take followed
One accepts either a script or an audio url and reads it aloud through one of five named providers. The other requires audio and says the recording normally determines the video length. Both put a face on screen. As of 2026-09-22.
| Field | D-ID | Hedra | Where they part |
|---|---|---|---|
| Audio source | From a script, or a supplied url | Supplied; audio is required | Different answers |
| Voice source | A voice id, or a recording by url | A track, or a voice id from the endpoint | Different answers |
| Per-character binding | On the request | Not documented by the vendor | Only D-ID answers |
| Languages | A language field, no list | A claim with no names | Different answers |
| Lip-sync | Not documented by the vendor | The supplied audio | Only Hedra answers |
Inclusion rule. Two entries are given a page together when at least one column puts them at opposite grades of answer. Pairs that agree on every column, or that are both blank throughout, do not get a page. Order. Fixed field order, identical on every side-by-side page.
1A required input and an optional one change the order of work
Where a recording is required, casting and a session sit in front of the first frame and a director approves the performance before anything is billed. Where a script is enough, the first output can exist within minutes and nobody has heard the line in advance.
Neither is wrong for every job. A team iterating on copy wants the script route; a team delivering a performance wants the recording. What matters is noticing which one a schedule was built around.
2Length as a parameter, and length as a property of the take
One publishes ceilings: five minutes for a clip, ten for a talk, forty thousand characters of script. The other says the audio normally determines the video length, which removes duration from the request and puts it in the recording.
Fitting a slot is therefore a different job on each. One is arithmetic against a published ceiling; the other is a re-record, which is a room and a performer rather than a number.
3Only one of them says what the mouth follows
The audio-driven entry states the character in the image lip-syncs and moves to the audio provided. The script-driven entry, whose endpoint is named for creating a talk, never mentions lip movement at all.
Both also claim language breadth without naming a language. One delegates to five providers, the other calls the support full, and neither leaves a reader able to check a market.
4Each of them on its own
The column this pair was chosen for is audio source, and each entry has a page of its own on it. The full row for either, all five columns with the wording behind each cell, is on its model note.
- D-ID on audio source — from a script, or a supplied url.
- Hedra on audio source — supplied; audio is required.
- D-ID, all five fields — read from the create a talk reference.
- Hedra, all five fields — read from the Character 3 model page and avatar video guide.
- Audio sourceA script is either text of up to forty thousand characters or an audio url, with audio limited to five minutes for clips and ten for talksa script read aloud or a file
- Voice sourceFive speech providers are named for the voice: Microsoft, ElevenLabs, Amazon, Google and Azure OpenAIoutside catalogues named on the page
- Audio sourceAvatar videos are driven by audio, and the character in the image will lip-sync and move to the audio providedhanded to the model rather than made by it
- Shot length set by the takeThe audio normally determines the video length, and omitting the duration follows the source audio
- LanguagesDescribed as text, image and audio to video with full multi-language support, and a maximum duration of ten minutesa claim carrying no list
5Sources
Every cell above is read from the documentation each vendor publishes, on the dates carried by the two model notes. The whole register on one column is on the field note; entries grouped by the shape of their answer are on the routes. Other pairs: Kling AI and Synthesia, SceneMixer and Synthesia.