Sentioscope

Speech and voice controls, as each vendor documents them

D-ID: a script read aloud, or a file handed over

The create-a-talk reference gives two ways in: a text script capped at forty thousand characters, or an audio url. Supplied audio runs to five minutes for a clip and ten for a talk, so the ceiling is written down rather than discovered. As of 2026-09-22.

Two ways into the same endpoint, and where the risk sitsA talk can be created from text of up to forty thousand characters or from an audio url. The text route fixes the words and leaves the delivery to a synthesiser; the audio route fixes the delivery and leaves the picture to follow.What is handed to the create-a-talk endpointText, up to 40,000 charactersA script to be readWords fixed, delivery decided by a voicenobody auditionedAn audio urlA recording to followFive minutes for a clip, ten for a talk,performance already approvedSame endpoint, opposite places to carry the risk
Fig. 1 Both routes reach the same call, so the choice is easy to make by accident and hard to read back off the returned file.
D-ID on audio source, statement by statement. Read from the vendor's create a talk reference on 2026-09-22.
What the documentation settlesWhat it leaves to a take
A script is either text of up to forty thousand characters, or an audio urlWhich of the two produces the steadier read of the same line
Supplied audio runs to five minutes for a clip, ten for a talkWhether a face holds attention across ten unbroken minutes
Five speech providers are named for the voiceWhich of the five a given voice id came from, before it is called

Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.

1Forty thousand characters is a ceiling nobody reaches

A forty-thousand-character script is roughly five hours of speech, which no single talking-head render is going to hold. Publishing the number anyway is useful in a different way: it says the text route is not a demo affordance with a hidden sentence limit, and a production can paste a whole scene into it without splitting the file first.

The audio route carries the ceiling that actually binds. Five minutes for a clip and ten for a talk are numbers a schedule can be built from, and they are the reason a long piece here is assembled from several calls rather than requested in one.

2Two routes that fail in opposite directions

Handing over a recording fixes the performance and leaves the picture to follow. Handing over text fixes the words and leaves the performance to a synthesiser that a director never auditioned. Neither is better in the abstract; they put the risk in different places, and only one of them can be approved before the render is paid for.

Because both routes land on the same endpoint, the choice is easy to make by accident. A team prototyping with text and delivering with recordings is testing something other than what it will ship, and nothing in the returned file says which route produced it.

3Others that make speech a step of its own

Three more entries here treat speech as a stage before the picture rather than as something the picture arrives with. What differs between them is how much of that stage is published.

  • HeyGen — from a speech endpoint, then rendered.
  • PixVerse — supplied, or read from text.
  • Synthesia — from the script, or uploaded.
  • Audio source
    A script is either text of up to forty thousand characters or an audio url, with audio limited to five minutes for clips and ten for talksa script read aloud or a fileD-ID, create a talk reference / recorded 2026-09-22
  • Voice source
    Five speech providers are named for the voice: Microsoft, ElevenLabs, Amazon, Google and Azure OpenAIoutside catalogues named on the pageD-ID, create a talk reference / recorded 2026-09-22

4Sources

Read from the create a talk reference at docs.d-id.com on 2026-09-22. The same column across every entry is on audio source; everything this vendor publishes about speech is on D-ID. What counts as documented is on how read.