What sync-3 documents about speech
sync-3 (sync.so) publishes a language count of ninety-five or more, accepts video or a still image paired with either audio or text, and matches lip movement to whatever track it is handed. As of 2026-09-22.
| Field | What the vendor documents |
|---|---|
| Audio source | Supplied, or read from text on the call |
| Languages | Ninety-five or more, counted rather than named |
| Voice source | Not documented on the model page |
| Per-character binding | Not documented |
| Lip-sync | Lip movement matched to the audio it is given |
Inclusion rule. Fields are the same five for every model here; a field the vendor does not document is recorded as not documented rather than inferred from generated output. Order. Fixed field order, identical on every model page.
1The largest language figure here, and still a figure
Ninety-five or more languages, described as the same coverage as the models before it, is the highest number in this register by a wide margin. It is also a count rather than a list, so it answers whether a language is likely to work and not whether a particular one does.
A count that high is plausible for a model working from a waveform, since there is no per-language voice inventory to build. That reasoning is not on the page, so it is not recorded as fact, and a production with one specific market still has to establish coverage the hard way.
2Four accepted pairs, and what each implies about the schedule
Video with audio, video with text, image with audio, image with text. The two audio pairs take a performance that already exists; the two text pairs have the speech produced on the way through. Choosing between them is choosing whether the actors record first.
The image pairs are the interesting ones for drama, because they turn a still into a talking shot without a camera having rolled. A free account runs one generation a month at fifteen seconds, which makes clear that evaluating this properly is a paid exercise.
3Alignment solved, casting elsewhere
Synchronisation is the product, so the lip-sync field fills itself. What stays empty is the voice: nothing on the model page says where one comes from, whether a choice persists, or what happens when two faces are in frame. A model attending to a waveform has no obvious way to know which mouth should move.
For a serial that leaves the same hole as everywhere else in this register, arriving from the opposite direction. The mouth is handled and the cast is somebody else's department.
4This entry, one column at a time
Each of these stays inside a single field: what sync-3 puts there, what the wording settles, and what it leaves for a take to answer.
- Audio source — supplied, or read from text.
- Voice source — nothing documented as an input.
- Languages — a count, ninety-five or more.
- Lip-sync — the audio it is given.
5Read against another entry
Each of these puts sync-3 beside one other entry on a column where the two land at opposite grades of answer.
- Vidu and sync-3 — on audio source.
- Synthesia and sync-3 — on languages.
6Where this model sits against the rest
| Model | Audio source |
|---|---|
| D-ID | A text script read aloud, or an audio url supplied |
| Hedra | Supplied; audio is a required input |
| HeyGen | A text to speech endpoint, apart from the picture |
| Kling AI | Generated with the picture, VIDEO 3.0 |
| LTX Studio | Generated jointly with the picture, plus audio-to-video |
| Luma Ray | Not documented; the video reference does not mention sound |
| MiniMax | Native speech, generated with the picture |
| PixVerse | Supplied to a speech endpoint, or read from text there |
| Runway | Not documented where voices are defined |
| SceneMixer | Generated with the picture, in the language set for the project |
| Sora 2 | Generated with the picture; video and audio are both listed as output |
| sync-3 | Supplied, or read from text on the call |
| Synthesia | Synthesised from the script, or an uploaded recording |
| Veo | Generated with the video rather than added afterwards |
| Vidu | Q3 generates speech, effects and music natively, on by default |
| Wan 3.0 | On by default; false returns a file with no audio track |
| Wan2.2-S2V | Supplied; the card describes the model as audio-driven |
Inclusion rule. Models whose vendor documents something about generated speech. A model with nothing published on the point is still listed, with the cell marked hollow. Order. Alphabetical by model name.
| Model | Languages |
|---|---|
| D-ID | A language field beside the voice; no list on the reference |
| Hedra | Full multi-language support claimed, none named |
| HeyGen | Thirty or more, counted on the translation endpoint |
| Kling AI | Five, not named |
| LTX Studio | Not documented |
| Luma Ray | Not documented |
| MiniMax | Not documented by the vendor |
| PixVerse | Multiple, with speech, singing and advertisements named as types |
| Runway | Not documented |
| SceneMixer | 15, named, with Cantonese a 16th for dialogue only |
| Sora 2 | Not documented |
| sync-3 | Ninety-five or more, counted rather than named |
| Synthesia | A catalogue, each voice carrying a language name and code |
| Veo | Not documented |
| Vidu | Not documented |
| Wan 3.0 | Not documented |
| Wan2.2-S2V | Not documented |
Inclusion rule. Models whose vendor documents something about generated speech. A model with nothing published on the point is still listed, with the cell marked hollow. Order. Alphabetical by model name.
| Model | Voice source |
|---|---|
| D-ID | A voice id from one of five named speech providers |
| Hedra | An uploaded track, or a voice id from the voices endpoint |
| HeyGen | A clone, instant from one recording or professional from 20 minutes or more |
| Kling AI | Not documented by the vendor |
| LTX Studio | Not documented |
| Luma Ray | Not documented |
| MiniMax | Reference audio, 15 s total across 3 clips |
| PixVerse | Built-in voices, or a custom voice from a supplied sample |
| Runway | An audio sample of 10 seconds to 5 minutes, or a description in words |
| SceneMixer | Preset voice library, or the user's own recordings |
| Sora 2 | Not documented |
| sync-3 | Not documented on the model page |
| Synthesia | The catalogue, or a cloned voice with a language list of its own |
| Veo | Not documented |
| Vidu | Not documented |
| Wan 3.0 | Reference audio, 15 seconds in total, WAV or MP3 up to 15 MB |
| Wan2.2-S2V | Whatever track is handed over; no catalogue |
Inclusion rule. Models whose vendor documents something about generated speech. A model with nothing published on the point is still listed, with the cell marked hollow. Order. Alphabetical by model name.
7Sources
Taken from the model documentation at sync.so on 2026-09-22. Every entry against the same five fields sits on the speech controls page; the rule behind each field is on how read. The entries either side of this one: Synthesia, Veo.
- Languagessync-3 supports 95 or more languages, described as the same coverage as the models before itpublished as a count
- Audio sourceThe accepted pairs are video with audio, video with text, image with audio and image with textsupplied or read from text
- Length per call on the free tierA free account runs one generation a month with a fifteen-second ceiling, with paid limits following the plan