sync-3: ninety-five languages and no voice at all
The accepted inputs include audio and text, and voices are never discussed. For a model that matches lip movement to whatever it is given, that is consistent: the voice is whoever made the recording, or whichever synthesiser read the text. As of 2026-09-22.
| What the documentation settles | What it leaves to a take |
|---|---|
| The accepted pairs are video with audio, video with text, image with audio and image with text | Which voice reads the text in the two text variants |
| sync-3 supports ninety-five or more languages | Whether a voice is available to a production in each of them |
| Lip movement is matched to the audio the model is given | Whether a synthesised read aligns as well as a recorded one |
Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.
1A text input implies a voice nobody documented
Two of the four accepted pairs take text rather than audio, which means something inside this product turns words into speech. That thing is a voice, and it is not described, counted, previewed or named. It is the clearest example in this register of a capability that exists by implication.
The register does not fill it in. Where a pipeline stage is undocumented, a production cannot plan a cast around it, and the honest cell is the one that says so rather than the one that assumes a reasonable default.
2The gap looks smaller once the product's position is clear
This is a stage rather than a generator: it operates on footage or a still that something else produced. Teams reaching it usually already have their audio, either from a session or from whichever speech service they standardised on, so a voice inventory here would go unused.
That explains the blank without removing it. A production intending to use the text variants is relying on an unnamed voice, and a season is a long time to rely on something no page describes.
3The others that publish nothing a production could hand over
Six more entries reach this column empty. Most of them make the sound themselves; this one accepts text and therefore has a voice it never mentions.
- Kling AI — nothing documented as an input.
- LTX Studio — nothing documented as an input.
- Luma Ray — nothing documented as an input.
- Sora 2 — nothing documented as an input.
- Veo — not documented by the vendor.
- Vidu — nothing documented as an input.
- Audio sourceThe accepted pairs are video with audio, video with text, image with audio and image with textsupplied or read from text
- Languagessync-3 supports 95 or more languages, described as the same coverage as the models before itpublished as a count
- Length per call on the free tierA free account runs one generation a month with a fifteen-second ceiling, with paid limits following the plan
4Sources
Read from the model documentation at sync.so on 2026-09-22. The same column across every entry is on voice source; everything this vendor publishes about speech is on sync-3. What counts as documented is on how read.