sync-3: four accepted pairs, and a free-tier ceiling
The model documentation lists four accepted pairs: video with audio, video with text, image with audio and image with text. Sound is therefore supplied or synthesised from text on the call, never produced alongside an invented picture. As of 2026-09-22.
| What the documentation settles | What it leaves to a take |
|---|---|
| The accepted pairs are video with audio, video with text, image with audio and image with text | Which pair gives the steadiest result for a dialogue scene |
| Lip movement is matched to the audio the model is given | How the match behaves when the supplied audio is noisy or overlapping |
| A free account runs one generation a month with a fifteen-second ceiling | What the paid ceilings are, since they are described as following the plan |
Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.
1Four pairs describe four different production shapes
Video with audio is a repair job on footage that already exists. Image with audio is a talking portrait. The two text variants move the speech synthesis inside this call instead of upstream of it. Listing all four makes the product's position explicit: it operates on material rather than inventing it.
For a drama pipeline that placement matters more than any parameter. This is a stage, and a stage has an input that something else must produce, which is a scheduling fact rather than a limitation.
2A published free ceiling is a real number in a field of vague ones
One generation a month with a fifteen-second ceiling is a small allowance and an unusually concrete one. It is enough to establish whether the alignment is acceptable on a production's own footage, which is exactly the question a page cannot answer.
The paid limits are described as following the plan rather than published as figures, so the register stops where the documentation does. A team that needs the ceiling in writing before committing has to ask for it.
3The other entries that are handed a recording
Two more entries take the sound as an input. Both of those are avatar or portrait products; this one also accepts existing footage, which puts it later in the pipeline.
- Hedra — supplied; audio is required.
- Wan2.2-S2V — supplied; the model is audio-driven.
- Audio sourceThe accepted pairs are video with audio, video with text, image with audio and image with textsupplied or read from text
- Length per call on the free tierA free account runs one generation a month with a fifteen-second ceiling, with paid limits following the plan
- Languagessync-3 supports 95 or more languages, described as the same coverage as the models before itpublished as a count
4Sources
Read from the model documentation at sync.so on 2026-09-22. The same column across every entry is on audio source; everything this vendor publishes about speech is on sync-3. What counts as documented is on how read.