Sentioscope

Speech and voice controls, as each vendor documents them

D-ID: a talk gets made, and the mouth goes unmentioned

The endpoint is called create a talk, and the reference never mentions lip movement. Scripts, audio ceilings and five voice providers are all published; what happens to the face while the words play is not. As of 2026-09-22.

An endpoint called create a talk, with no mouth on itScripts, audio ceilings and five voice providers are all published. The word for the thing the product exists to do does not appear, which makes this the most striking blank in the column.Audio suppliedScript suppliedWhat alignment followsA waveform that already existsSpeech synthesised first, then followedExtra stepsNoneOne, and one more place to lose a framePublished ceilingFive minutes, ten for a talkForty thousand charactersStated qualityNothingNothingInterchangeable on the request, possibly not in the result
Fig. 1 The two routes into the endpoint may not align equally well, and a team will drift between them without noticing.
D-ID on lip-sync, statement by statement. Read from the vendor's create a talk reference on 2026-09-22.
What the documentation settlesWhat it leaves to a take
A script is either text of up to forty thousand characters or an audio urlWhether the mouth follows the audio or the text it came from
Five speech providers are named for the voiceWhether alignment differs between provider voices
Nothing on the reference addresses lip movementHow the face behaves across a ten-minute talk

Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.

1A silence that is harder to excuse than most here

Pages written about picture generation can be forgiven for skipping mouths. This one is written about producing a talking person, and the word for the thing the product exists to do does not appear on it. That is the most striking blank in this column.

The likely explanation is that the reference is an API surface rather than a product description: it lists fields a caller sets, and there is no field for a convincing mouth. That explains the omission without making it less consequential for a reader choosing between tools.

2The audio route and the text route may not behave alike

Where a recording is supplied, alignment has a waveform to follow. Where a script is supplied, speech is synthesised first and then presumably followed, which is an extra step and an extra place to lose a frame. Nothing says whether the two produce the same quality.

For a production that difference is testable and worth testing, because the two routes are otherwise interchangeable on this endpoint and a team will drift between them without noticing.

3Others whose pages never raise the subject

Six more entries reach this column with nothing in it. Most of those documents are about generating pictures; this one is about generating a person who talks.

  • LTX Studio — not documented by the vendor.
  • Luma Ray — not documented by the vendor.
  • Runway — not documented by the vendor.
  • Veo — not documented by the vendor.
  • Vidu — not documented by the vendor.
  • Wan 3.0 — not documented by the vendor.
  • Audio source
    A script is either text of up to forty thousand characters or an audio url, with audio limited to five minutes for clips and ten for talksa script read aloud or a fileD-ID, create a talk reference / recorded 2026-09-22
  • Voice source
    Five speech providers are named for the voice: Microsoft, ElevenLabs, Amazon, Google and Azure OpenAIoutside catalogues named on the pageD-ID, create a talk reference / recorded 2026-09-22
  • Per-character binding
    The voice is a voice id selected from the list of available voices, with an optional language field beside ita voice id on every requestD-ID, create a talk reference / recorded 2026-09-22

4Sources

Read from the create a talk reference at docs.d-id.com on 2026-09-22. The same column across every entry is on lip-sync; everything this vendor publishes about speech is on D-ID. What counts as documented is on how read.