Sentioscope

Speech and voice controls, as each vendor documents them

Entries that answer with a number and no names

Three entries answer with a number. Five, thirty or more, and ninety-five or more. The largest belongs to a model that works from a waveform and so has no per-language voice inventory to publish in the first place. As of 2026-09-22.

Three totals that are not measuring the same thingA count attached to native audio on a model version, a count attached to video translation, and a count describing alignment across phonetics. None of the three transfers to either of the others.Where each total is attachedFive, on a model versionNative audio on one line, so a cheaper line may not carry it.Thirty or more, ontranslationReached by translating a finished master, not by choosing first.Ninety-five or more, onalignmentNo voice per language; the waveform arrives from elsewhere.And why none of them is comparable to the others
Fig. 1 A reader comparing five against ninety-five side by side is comparing different quantities.
The three published totals, beside the architecture each one describes. Recorded 2026-09-22.
ModelThe published totalAnd where the sound is made
HeyGenA count, attached to translationFrom a speech endpoint, then rendered
Kling AIA count, no namesWith the picture, on VIDEO 3.0
sync-3A count, ninety-five or moreSupplied, or read from text

Inclusion rule. Entries whose documentation gives a number of spoken languages without naming them. An unnamed claim of breadth with no number is on a separate route page. Order. Alphabetical by model name.

1A large total can mean less work rather than more

Ninety-five languages sounds like an enormous inventory and probably describes the opposite. A model animating a mouth to supplied audio needs no voice per language; it needs alignment to hold across the sounds those languages make. That is a claim about phonetics, not about casting.

Read that way the figure becomes more credible and less useful. Credible because alignment generalises; less useful because it says nothing about whether a production can obtain a voice in the language it needs.

2Where a total sits tells you what it is for

One of these attaches its figure to video translation, which describes a workflow: one master, one performance, and a localisation stage at the end. Another attaches it to native audio on a specific model version, which makes the capability a property of a version rather than of an account.

Neither figure transfers to the other arrangement. A count that belongs to translation says nothing about generating in a language from the start, and a reader comparing the two numbers side by side is comparing different things.

3The smallest total is the one that hurts most

With a large number a production can assume its market is probably included and check cheaply. With five, the probability runs the other way, and an assumption becomes a commitment made on a coin toss.

Five is also short enough that naming the five would have been decisive, which is why the omission reads less like an editorial decision about page length and more like a capability still settling.

4The entries on this route, one page each

Each of these links to that entry on the field this route groups by. The wording behind every cell, and the date it was read, sits on the page it links to.

  • HeyGen — a count, attached to translation.
  • Kling AI — a count, no names.
  • sync-3 — a count, ninety-five or more.
  • Languages
    sync-3 supports 95 or more languages, described as the same coverage as the models before itpublished as a countSync, sync-3 model documentation / recorded 2026-09-22
  • Languages
    Video translation is documented for thirty or more languages, with voice cloning and lip-synccounted on the translation endpointHeyGen, API quick start / recorded 2026-09-22
  • Audio source
    Native audio on VIDEO 3.0 in five languagesgenerated with the pictureKling AI, model guide / recorded 2026-09-12
  • Audio source
    The accepted pairs are video with audio, video with text, image with audio and image with textsupplied or read from textSync, sync-3 model documentation / recorded 2026-09-22

5Sources

Membership of this route is decided by the wording in the field it groups by, read from the documentation each vendor publishes. The column itself is on the field note, all five columns are on the speech table, and what counts as documented is on how read. Other routes: Breadth claimed, A figure in writing.