Sentioscope

Speech and voice controls, as each vendor documents them

sync-3: ninety-five languages and no voice at all

The accepted inputs include audio and text, and voices are never discussed. For a model that matches lip movement to whatever it is given, that is consistent: the voice is whoever made the recording, or whichever synthesiser read the text. As of 2026-09-22.

A text input implies a voice the page never namesTwo of the four accepted pairs take text rather than audio, so something inside the product turns words into speech. That thing is a voice, and it is not described, counted, previewed or named anywhere.Audio inText inWho speaksWhoever made the recordingSomething the page never namesWhat is documentedThe pairs, and that lips follow audioThe pairs, and nothing elseCan it be auditionedYes, before the callNot from anything publishedWhere a cast livesIn a recording sessionNowhere statedThe largest language count here arrives with no casting mechanism
Fig. 1 Where a pipeline stage is undocumented a production cannot plan a cast around it, so the honest cell says so rather than assuming a default.
sync-3 on voice source, statement by statement. Read from the vendor's model documentation on 2026-09-22.
What the documentation settlesWhat it leaves to a take
The accepted pairs are video with audio, video with text, image with audio and image with textWhich voice reads the text in the two text variants
sync-3 supports ninety-five or more languagesWhether a voice is available to a production in each of them
Lip movement is matched to the audio the model is givenWhether a synthesised read aligns as well as a recorded one

Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.

1A text input implies a voice nobody documented

Two of the four accepted pairs take text rather than audio, which means something inside this product turns words into speech. That thing is a voice, and it is not described, counted, previewed or named. It is the clearest example in this register of a capability that exists by implication.

The register does not fill it in. Where a pipeline stage is undocumented, a production cannot plan a cast around it, and the honest cell is the one that says so rather than the one that assumes a reasonable default.

2The gap looks smaller once the product's position is clear

This is a stage rather than a generator: it operates on footage or a still that something else produced. Teams reaching it usually already have their audio, either from a session or from whichever speech service they standardised on, so a voice inventory here would go unused.

That explains the blank without removing it. A production intending to use the text variants is relying on an unnamed voice, and a season is a long time to rely on something no page describes.

3The others that publish nothing a production could hand over

Six more entries reach this column empty. Most of them make the sound themselves; this one accepts text and therefore has a voice it never mentions.

  • Kling AI — nothing documented as an input.
  • LTX Studio — nothing documented as an input.
  • Luma Ray — nothing documented as an input.
  • Sora 2 — nothing documented as an input.
  • Veo — not documented by the vendor.
  • Vidu — nothing documented as an input.
  • Audio source
    The accepted pairs are video with audio, video with text, image with audio and image with textsupplied or read from textSync, sync-3 model documentation / recorded 2026-09-22
  • Languages
    sync-3 supports 95 or more languages, described as the same coverage as the models before itpublished as a countSync, sync-3 model documentation / recorded 2026-09-22
  • Length per call on the free tier
    A free account runs one generation a month with a fifteen-second ceiling, with paid limits following the planSync, sync-3 model documentation / recorded 2026-09-22

4Sources

Read from the model documentation at sync.so on 2026-09-22. The same column across every entry is on voice source; everything this vendor publishes about speech is on sync-3. What counts as documented is on how read.