Sentioscope

Speech and voice controls, as each vendor documents them

D-ID and Hedra: a script read out, or a take followed

One accepts either a script or an audio url and reads it aloud through one of five named providers. The other requires audio and says the recording normally determines the video length. Both put a face on screen. As of 2026-09-22.

A script read out, against a take that is followedWhere a recording is required, casting and a session sit in front of the first frame. Where a script is enough, the first output exists within minutes and nobody has heard the line in advance.How the words reach the screenA script, up to 40,000 charactersRead aloud for youFast, and nobody auditioned the deliveryA recording you supplyFollowed as givenSlower at the front, and the performanceis approvedLength is a published ceiling on one and a property of the take on the other
Fig. 1 What matters is noticing which of the two a schedule was built around.
Controls documented, model by modelFilled where the model documents that control, hollow where nothing is published about it.Controls documented, model by modelD-IDHedraWhere they partAudio sourceAudio source — D-ID: From a script, or a supplied urlAudio source — Hedra: Supplied; audio is requiredAudio source — Where they part: Different answersVoice sourceVoice source — D-ID: A voice id, or a recording by urlVoice source — Hedra: A track, or a voice id from the endpointVoice source — Where they part: Different answersPer-character bindingPer-character binding — D-ID: On the requestPer-character binding — Hedra: Not documented by the vendorPer-character binding — Where they part: Only D-ID answersLanguagesLanguages — D-ID: A language field, no listLanguages — Hedra: A claim with no namesLanguages — Where they part: Different answersLip-syncLip-sync — D-ID: Not documented by the vendorLip-sync — Hedra: The supplied audioLip-sync — Where they part: Only Hedra answers
Fig. 2 Filled where the model documents that control, hollow where nothing is published about it.
Models documenting each controlHow many models document each control. A hollow column is a statement about documentation, not capability.Models documenting each controlD-ID4 of 5Hedra4 of 5Where they part5 of 5
Fig. 3 How many models document each control. A hollow column is a statement about documentation, not capability.
D-ID and Hedra on the five fields this register keeps, with the gap on each. Recorded 2026-09-22.
FieldD-IDHedraWhere they part
Audio sourceFrom a script, or a supplied urlSupplied; audio is requiredDifferent answers
Voice sourceA voice id, or a recording by urlA track, or a voice id from the endpointDifferent answers
Per-character bindingOn the requestNot documented by the vendorOnly D-ID answers
LanguagesA language field, no listA claim with no namesDifferent answers
Lip-syncNot documented by the vendorThe supplied audioOnly Hedra answers

Inclusion rule. Two entries are given a page together when at least one column puts them at opposite grades of answer. Pairs that agree on every column, or that are both blank throughout, do not get a page. Order. Fixed field order, identical on every side-by-side page.

1A required input and an optional one change the order of work

Where a recording is required, casting and a session sit in front of the first frame and a director approves the performance before anything is billed. Where a script is enough, the first output can exist within minutes and nobody has heard the line in advance.

Neither is wrong for every job. A team iterating on copy wants the script route; a team delivering a performance wants the recording. What matters is noticing which one a schedule was built around.

2Length as a parameter, and length as a property of the take

One publishes ceilings: five minutes for a clip, ten for a talk, forty thousand characters of script. The other says the audio normally determines the video length, which removes duration from the request and puts it in the recording.

Fitting a slot is therefore a different job on each. One is arithmetic against a published ceiling; the other is a re-record, which is a room and a performer rather than a number.

3Only one of them says what the mouth follows

The audio-driven entry states the character in the image lip-syncs and moves to the audio provided. The script-driven entry, whose endpoint is named for creating a talk, never mentions lip movement at all.

Both also claim language breadth without naming a language. One delegates to five providers, the other calls the support full, and neither leaves a reader able to check a market.

4Each of them on its own

The column this pair was chosen for is audio source, and each entry has a page of its own on it. The full row for either, all five columns with the wording behind each cell, is on its model note.

  • Audio source
    A script is either text of up to forty thousand characters or an audio url, with audio limited to five minutes for clips and ten for talksa script read aloud or a fileD-ID, create a talk reference / recorded 2026-09-22
  • Voice source
    Five speech providers are named for the voice: Microsoft, ElevenLabs, Amazon, Google and Azure OpenAIoutside catalogues named on the pageD-ID, create a talk reference / recorded 2026-09-22
  • Audio source
    Avatar videos are driven by audio, and the character in the image will lip-sync and move to the audio providedhanded to the model rather than made by itHedra, avatar video guide / recorded 2026-09-22
  • Shot length set by the take
    The audio normally determines the video length, and omitting the duration follows the source audioHedra, avatar video guide / recorded 2026-09-22
  • Languages
    Described as text, image and audio to video with full multi-language support, and a maximum duration of ten minutesa claim carrying no listHedra, Character 3 model page / recorded 2026-09-22

5Sources

Every cell above is read from the documentation each vendor publishes, on the dates carried by the two model notes. The whole register on one column is on the field note; entries grouped by the shape of their answer are on the routes. Other pairs: Kling AI and Synthesia, SceneMixer and Synthesia.