Sentioscope

Speech and voice controls, as each vendor documents them

sync-3: ninety-five or more, and still only a number

Ninety-five or more is the largest figure in this column, and it belongs to a model that works from a waveform. There is no per-language voice inventory to publish, because the voice was recorded by whoever made the audio. As of 2026-09-22.

A total about alignment, not about castingA model that animates a mouth to a supplied waveform needs no voice for each language. Ninety-five or more is therefore a claim about phonetic coverage, which makes it more credible and less useful than it looks.A count of voicesA count of coverageWhat has to exist per languageA voice, cast and recordedNothing; the waveform arrivesWhat the number describesAn inventory somebody builtAlignment holding across soundsWhy it can be largeSustained investmentBecause it generalisesWhat it tells a productionWhether a voice is availableWhether a supplied voice will landThe largest total here is the second kind
Fig. 1 Describing the count as unchanged from earlier models dates the figure, which is more than most totals in this column offer.
sync-3 on languages, statement by statement. Read from the vendor's model documentation on 2026-09-22.
What the documentation settlesWhat it leaves to a take
sync-3 supports ninety-five or more languages, the same coverage as the models before itWhich languages, since a total is all that appears
The accepted pairs are video with audio, video with text, image with audio and image with textWhich of those pairs the language count applies to
Lip movement is matched to the audio the model is givenHow accurately the match holds in a language with unusual mouth shapes

Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.

1A large total can mean less work rather than more

Ninety-five languages sounds like an enormous inventory and probably describes the opposite. A model that animates a mouth to a supplied waveform does not need a voice for each language; it needs alignment to hold across the sounds those languages make. The count is a claim about phonetic coverage, not about casting.

Read that way the figure becomes more credible and less useful. Credible because alignment generalises; less useful because it says nothing about whether a production can obtain a voice in the language it needs.

2Same coverage as the models before it is a version statement

Describing the count as unchanged from earlier models is an unusual thing to publish and a helpful one. It tells a reader the language reach is not what this version improved, which narrows down what a migration would gain and, more practically, means a production on an older version has nothing to chase here.

It also quietly dates the figure. A number carried forward across versions has been true for some time, which is more than can be said for most totals in this column.

3The others that publish a total with no names

Two more entries answer with a number. Both of those attach the figure to a voice inventory of some kind, which is what makes this one different: the total describes alignment rather than casting.

  • HeyGen — a count, attached to translation.
  • Kling AI — a count, no names.
  • Languages
    sync-3 supports 95 or more languages, described as the same coverage as the models before itpublished as a countSync, sync-3 model documentation / recorded 2026-09-22
  • Audio source
    The accepted pairs are video with audio, video with text, image with audio and image with textsupplied or read from textSync, sync-3 model documentation / recorded 2026-09-22

4Sources

Read from the model documentation at sync.so on 2026-09-22. The same column across every entry is on languages; everything this vendor publishes about speech is on sync-3. What counts as documented is on how read.