sync-3: movement matched to the audio it is given
Lip movement is matched to the audio the model is given. Where the waveform is an input rather than an output, naming the driver costs the vendor nothing and tells a production exactly which half it controls. As of 2026-09-22.
| What the documentation settles | What it leaves to a take |
|---|---|
| Lip movement is matched to the audio the model is given | How the match behaves when the audio is noisy or overlapping |
| The accepted pairs are video with audio, video with text, image with audio and image with text | Whether the text variants align as well as the audio ones |
| A free account runs one generation a month with a fifteen-second ceiling | What the alignment looks like on a production's own footage |
Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.
1A product whose whole job is this column
Most entries here mention mouths in passing. This one exists to move them, so the driver statement is not a concession but a description of the product. The interesting questions move downstream: how well, on what material, and at what cost.
The free tier is the practical answer to the first of those. One generation a month with a fifteen-second ceiling is enough to try the alignment on real footage, which is exactly the test a page cannot substitute for.
2The text variants are a quieter claim
Two of the four accepted pairs take text, which means speech is synthesised somewhere before the mouth is matched. The driver statement is about audio, so a reader has to assume the synthesised audio plays the same role, and the page does not confirm it.
Recorded as a driver named for the audio routes. Whether the text routes behave identically is testable and unstated, and it matters because those routes are the ones a team with no recording session will reach for.
3The others that name what the mouth is following
Four more entries name a driver. Two are audio-driven like this one, one drives the mouth from a script and attaches a framing condition, and one makes the sound itself.
- Hedra — the supplied audio.
- MiniMax — whoever is on screen.
- Synthesia — the spoken content, framed close.
- Wan2.2-S2V — the audio input.
- Audio sourceThe accepted pairs are video with audio, video with text, image with audio and image with textsupplied or read from text
- Length per call on the free tierA free account runs one generation a month with a fifteen-second ceiling, with paid limits following the plan
- Languagessync-3 supports 95 or more languages, described as the same coverage as the models before itpublished as a count
4Sources
Read from the model documentation at sync.so on 2026-09-22. The same column across every entry is on lip-sync; everything this vendor publishes about speech is on sync-3. What counts as documented is on how read.