Synthesia and D-ID: a driver named, or never mentioned
Both products exist to put a talking person on screen. One documents lip sync generated from the spoken content, with a framing condition attached. The other creates a talk and never mentions the mouth. As of 2026-09-22.
| Field | Synthesia | D-ID | Where they part |
|---|---|---|---|
| Audio source | From the script, or uploaded | From a script, or a supplied url | Different answers |
| Voice source | A catalogue voice, or a cloned one | A voice id, or a recording by url | Different answers |
| Per-character binding | Not documented by the vendor | On the request | Only D-ID answers |
| Languages | A catalogue with codes | A language field, no list | Different answers |
| Lip-sync | The spoken content, framed close | Not documented by the vendor | Only Synthesia answers |
Inclusion rule. Two entries are given a page together when at least one column puts them at opposite grades of answer. Pairs that agree on every column, or that are both blank throughout, do not get a page. Order. Fixed field order, identical on every side-by-side page.
1The starkest pair of blanks in this register
An endpoint literally called create a talk that never says what happens to a mouth is the most striking omission the lip-sync column found. Set beside a competitor that names both a driver and a condition, it reads as an API surface rather than a product description.
That is an explanation and not an excuse. A producer choosing between the two before any money moves has only the documents, and one of them has skipped the part an audience notices first.
2Their voice arrangements are opposite too
One publishes its own catalogue, with a language name, a code, a gender and an id on every voice, plus a cloning route with its own range. The other names five outside speech providers and leaves the enumeration to them.
The first can be shortlisted before anything is signed. The second requires auditing five upstream catalogues, each with its own licence terms, deprecation schedule and language codes, none of which reaches a reader through the talk reference.
3Both publish real figures, and about different things
One publishes a threshold on nothing and a condition on framing; the other publishes forty thousand characters of script and audio ceilings of five and ten minutes. Both are numbers a schedule can use, and they bound completely different quantities.
Lined up, they show why the published-figure route page groups by the existence of a number rather than by what it measures. Vendors are specific about very different things.
4Each of them on its own
The column this pair was chosen for is lip-sync, and each entry has a page of its own on it. The full row for either, all five columns with the wording behind each cell, is on its model note.
- Synthesia on lip-sync — the spoken content, framed close.
- D-ID on lip-sync — not documented by the vendor.
- Synthesia, all five fields — read from the voices reference and avatar documentation.
- D-ID, all five fields — read from the create a talk reference.
- Lip-syncLip sync and facial expressions are generated from the spoken contentdriven by the script
- A framing condition attached to lip-syncLip sync performs best when the avatar is framed closer in the scene rather than positioned far from the camera
- LanguagesEach voice is listed with its formal and native language name, a language code, a gender, a name and a voice idpublished as a catalogue with codes
- Audio sourceA script is either text of up to forty thousand characters or an audio url, with audio limited to five minutes for clips and ten for talksa script read aloud or a file
- Voice sourceFive speech providers are named for the voice: Microsoft, ElevenLabs, Amazon, Google and Azure OpenAIoutside catalogues named on the page
5Sources
Every cell above is read from the documentation each vendor publishes, on the dates carried by the two model notes. The whole register on one column is on the field note; entries grouped by the shape of their answer are on the routes. Other pairs: HeyGen and Synthesia, D-ID and Hedra.