Sentioscope

Speech and voice controls, as each vendor documents them

Entries whose pages never raise the mouth at all

Seven entries never raise the subject on the page that was read. Several of them plainly put a talking person on screen, so the blank belongs to the document rather than to the product and it has a consistent shape. As of 2026-09-22.

Why a parameter list skips the visible halfMost of these pages are API references. They describe what a caller sets, and nothing a caller sets makes a mouth look right, so the subject falls outside what the document is for. Two of the seven have no such excuse.Pages with an excusePages without oneWhat they areVideo references and picture APIsA talking-head endpoint, and a voicesreferenceWhy the blank fitsNo parameter makes a mouth convincingThe document is about speechWhat they do settleWhere sound comes from, sometimesScripts, ceilings, samples, previewsHow the cell readsUnstated, and understandableUnstated, and the most striking gap hereSilence is not evidence that the case is handled
Fig. 1 Filling these cells from plausibility would reward whichever vendor a reader already trusted.
Controls documented, model by modelFilled where the model documents that control, hollow where nothing is published about it.Controls documented, model by modelWhat the page saysAnd what it does addressAnd what it doesaddressD-IDD-ID — What the page says: Not documented by the vendorD-ID — And what it does address: From a script, or a supplied urlLTX StudioLTX Studio — What the page says: Not documented by the vendorLTX Studio — And what it does address: With the picture, and audio to videoLuma RayLuma Ray — What the page says: Not documented by the vendorLuma Ray — And what it does address: Not documented by the vendorRunwayRunway — What the page says: Not documented by the vendorRunway — And what it does address: Not documented by the vendorVeoVeo — What the page says: Not documented by the vendorVeo — And what it does address: With the pictureViduVidu — What the page says: Not documented by the vendorVidu — And what it does address: With the picture, speech available on its ownWan 3.0Wan 3.0 — What the page says: Not documented by the vendorWan 3.0 — And what it does address: With the picture, unless switched off
Fig. 2 Filled where the model documents that control, hollow where nothing is published about it.
Entries silent on the mouth, beside what their pages do settle about sound. Recorded 2026-09-22.
ModelWhat the page saysAnd what it does address
D-IDNot documented by the vendorFrom a script, or a supplied url
LTX StudioNot documented by the vendorWith the picture, and audio to video
Luma RayNot documented by the vendorNot documented by the vendor
RunwayNot documented by the vendorNot documented by the vendor
VeoNot documented by the vendorWith the picture
ViduNot documented by the vendorWith the picture, speech available on its own
Wan 3.0Not documented by the vendorWith the picture, unless switched off

Inclusion rule. Entries whose documentation, as read for this register, says nothing about the relationship between a generated mouth and speech. An entry naming lip-sync without a driver is on a separate route page. Order. Alphabetical by model name.

1A parameter list has no field for a convincing mouth

Most of these pages are API references. They describe what a caller sets, and nothing a caller sets makes a mouth look right, so the subject falls outside what the document is for. That explains the blank without excusing it.

It also predicts where the exceptions come from. The entries that do name a driver are mostly products where the audio is an input, because there the mouth's behaviour is a property of the interface and belongs on a reference.

2Two of these blanks are much harder to excuse

One entry here is an endpoint for creating a talking person, and the word for what it does to a mouth is not on the page. Another is the most thorough voice-construction reference in the register and stops before any performance.

Those two are not documents that ran out of room on an unrelated subject. They are documents about speech that leave out the visible half of it, which is the most striking pattern this column found.

3Silence is not evidence that the case is handled

Several products in this group generate speech with the picture, and something in them is aligning a mouth to it. A production can reasonably infer that the joint generation handles it, and an inference is worth less than a sentence when a shot comes back wrong.

Keeping these cells empty is what lets the rows be compared. Filling them from plausibility would reward whichever vendor a reader already trusted, which is the failure mode this whole register is built to avoid.

4The entries on this route, one page each

Each of these links to that entry on the field this route groups by. The wording behind every cell, and the date it was read, sits on the page it links to.

  • D-ID — not documented by the vendor.
  • LTX Studio — not documented by the vendor.
  • Luma Ray — not documented by the vendor.
  • Runway — not documented by the vendor.
  • Veo — not documented by the vendor.
  • Vidu — not documented by the vendor.
  • Wan 3.0 — not documented by the vendor.
  • Audio source
    Q3 outputs speech, sound effects and music natively, on by default in the API, with a speech-only modethree tracks in one passVidu, API reference / recorded 2026-09-12
  • Audio source
    Audio is generated with the video rather than added afterwardsgenerated with the pictureGoogle, Veo documentation / recorded 2026-09-12
  • Audio source
    The audio parameter defaults to true so the returned video carries an audio track, and false returns one withouton unless it is switched offAlibaba Cloud, Wan 3.0 API reference / recorded 2026-09-22
  • Audio source
    Joint audio and video generation, plus audio-to-video, on LTX-2.5two directionsLTX Studio / recorded 2026-09-12
  • Audio source
    Not documented by the vendor (as of 2026-09-22)no mention of sound in the video referenceLuma, video generation reference / recorded 2026-09-22
  • Audio source
    A script is either text of up to forty thousand characters or an audio url, with audio limited to five minutes for clips and ten for talksa script read aloud or a fileD-ID, create a talk reference / recorded 2026-09-22
  • Lip-sync
    Not documented by the vendor (as of 2026-09-22)not addressed where voices are definedRunway, custom voices reference / recorded 2026-09-22

5Sources

Membership of this route is decided by the wording in the field it groups by, read from the documentation each vendor publishes. The column itself is on the field note, all five columns are on the speech table, and what counts as documented is on how read. Other routes: Sound in the same pass, Handed a recording.