Sentioscope

Speech and voice controls, as each vendor documents them

Synthesia and D-ID: a driver named, or never mentioned

Both products exist to put a talking person on screen. One documents lip sync generated from the spoken content, with a framing condition attached. The other creates a talk and never mentions the mouth. As of 2026-09-22.

Two talking-head products, one silent about the mouthAn endpoint called create a talk that never says what happens to a mouth, beside a competitor that names both a driver and a framing condition. The starkest pair of blanks this column found.SynthesiaD-IDThe mouthFrom the spoken content, framed closeNever mentionedThe voicesIts own catalogue, with codes and idsFive outside providers, unenumeratedAuditing the voicesShortlist before signing anythingAudit five upstream cataloguesPublished figuresA framing condition40,000 characters; five and ten minutesBoth publish real numbers, about different things
Fig. 1 A producer choosing between the two before any money moves has only the documents, and one has skipped the part an audience notices first.
Controls documented, model by modelFilled where the model documents that control, hollow where nothing is published about it.Controls documented, model by modelSynthesiaD-IDWhere they partAudio sourceAudio source — Synthesia: From the script, or uploadedAudio source — D-ID: From a script, or a supplied urlAudio source — Where they part: Different answersVoice sourceVoice source — Synthesia: A catalogue voice, or a cloned oneVoice source — D-ID: A voice id, or a recording by urlVoice source — Where they part: Different answersPer-character bindingPer-character binding — Synthesia: Not documented by the vendorPer-character binding — D-ID: On the requestPer-character binding — Where they part: Only D-ID answersLanguagesLanguages — Synthesia: A catalogue with codesLanguages — D-ID: A language field, no listLanguages — Where they part: Different answersLip-syncLip-sync — Synthesia: The spoken content, framed closeLip-sync — D-ID: Not documented by the vendorLip-sync — Where they part: Only Synthesia answers
Fig. 2 Filled where the model documents that control, hollow where nothing is published about it.
Models documenting each controlHow many models document each control. A hollow column is a statement about documentation, not capability.Models documenting each controlSynthesia4 of 5D-ID4 of 5Where they part5 of 5
Fig. 3 How many models document each control. A hollow column is a statement about documentation, not capability.
Synthesia and D-ID on the five fields this register keeps, with the gap on each. Recorded 2026-09-22.
FieldSynthesiaD-IDWhere they part
Audio sourceFrom the script, or uploadedFrom a script, or a supplied urlDifferent answers
Voice sourceA catalogue voice, or a cloned oneA voice id, or a recording by urlDifferent answers
Per-character bindingNot documented by the vendorOn the requestOnly D-ID answers
LanguagesA catalogue with codesA language field, no listDifferent answers
Lip-syncThe spoken content, framed closeNot documented by the vendorOnly Synthesia answers

Inclusion rule. Two entries are given a page together when at least one column puts them at opposite grades of answer. Pairs that agree on every column, or that are both blank throughout, do not get a page. Order. Fixed field order, identical on every side-by-side page.

1The starkest pair of blanks in this register

An endpoint literally called create a talk that never says what happens to a mouth is the most striking omission the lip-sync column found. Set beside a competitor that names both a driver and a condition, it reads as an API surface rather than a product description.

That is an explanation and not an excuse. A producer choosing between the two before any money moves has only the documents, and one of them has skipped the part an audience notices first.

2Their voice arrangements are opposite too

One publishes its own catalogue, with a language name, a code, a gender and an id on every voice, plus a cloning route with its own range. The other names five outside speech providers and leaves the enumeration to them.

The first can be shortlisted before anything is signed. The second requires auditing five upstream catalogues, each with its own licence terms, deprecation schedule and language codes, none of which reaches a reader through the talk reference.

3Both publish real figures, and about different things

One publishes a threshold on nothing and a condition on framing; the other publishes forty thousand characters of script and audio ceilings of five and ten minutes. Both are numbers a schedule can use, and they bound completely different quantities.

Lined up, they show why the published-figure route page groups by the existence of a number rather than by what it measures. Vendors are specific about very different things.

4Each of them on its own

The column this pair was chosen for is lip-sync, and each entry has a page of its own on it. The full row for either, all five columns with the wording behind each cell, is on its model note.

  • Lip-sync
    Lip sync and facial expressions are generated from the spoken contentdriven by the scriptSynthesia, create an avatar / recorded 2026-09-22
  • A framing condition attached to lip-sync
    Lip sync performs best when the avatar is framed closer in the scene rather than positioned far from the cameraSynthesia, create an avatar / recorded 2026-09-22
  • Languages
    Each voice is listed with its formal and native language name, a language code, a gender, a name and a voice idpublished as a catalogue with codesSynthesia, list of supported voices / recorded 2026-09-22
  • Audio source
    A script is either text of up to forty thousand characters or an audio url, with audio limited to five minutes for clips and ten for talksa script read aloud or a fileD-ID, create a talk reference / recorded 2026-09-22
  • Voice source
    Five speech providers are named for the voice: Microsoft, ElevenLabs, Amazon, Google and Azure OpenAIoutside catalogues named on the pageD-ID, create a talk reference / recorded 2026-09-22

5Sources

Every cell above is read from the documentation each vendor publishes, on the dates carried by the two model notes. The whole register on one column is on the field note; entries grouped by the shape of their answer are on the routes. Other pairs: HeyGen and Synthesia, D-ID and Hedra.