Vidu and sync-3: three tracks made, or one matched
Vidu documents three audio tracks generated natively and on by default, with a speech-only mode beside them. sync-3 matches lip movement to audio it is given. One invents a soundscape; the other repairs a performance. As of 2026-09-22.
| Field | Vidu | sync-3 | Where they part |
|---|---|---|---|
| Audio source | With the picture, speech available on its own | Supplied, or read from text | Different answers |
| Voice source | Nothing documented as an input | Nothing documented as an input | Same answer |
| Per-character binding | Not documented by the vendor | Not documented by the vendor | Same answer |
| Languages | Not documented by the vendor | A count, ninety-five or more | Only sync-3 answers |
| Lip-sync | Not documented by the vendor | The audio it is given | Only sync-3 answers |
Inclusion rule. Two entries are given a page together when at least one column puts them at opposite grades of answer. Pairs that agree on every column, or that are both blank throughout, do not get a page. Order. Fixed field order, identical on every side-by-side page.
1A default soundscape against a supplied one
One of these returns a finished mix unless the call says otherwise, which makes silence something a production has to ask for. The other returns nothing until a track is handed over, so silence is the default and sound is the input.
Neither is better and they belong at different points in a pipeline. The first is a generator whose output is nearly a deliverable; the second is a stage applied to something already made.
2The speech-only mode is the interesting overlap
Keeping the dialogue and dropping the bed is documented on only one entry in this register, and it is the generator. That matters because the usual objection to generated audio is the music, not the speech, and a mode that separates them removes the objection without removing the feature.
The other entry never has that problem, because the bed was never generated. What it has instead is a dependency: something upstream must produce the dialogue, and that something is not described on its page.
3One names a driver, and the other names nothing about mouths
Matching lip movement to supplied audio is stated plainly on one page. The other describes three tracks and never says what happens to a face, which is the pattern for native-audio entries throughout this register.
Set side by side that is the clearest argument for keeping the lip-sync column at all. Two products that both put speech on screen give a reader completely different amounts to plan with.
4Each of them on its own
The column this pair was chosen for is audio source, and each entry has a page of its own on it. The full row for either, all five columns with the wording behind each cell, is on its model note.
- Vidu on audio source — with the picture, speech available on its own.
- sync-3 on audio source — supplied, or read from text.
- Vidu, all five fields — read from the API reference.
- sync-3, all five fields — read from the model documentation.
- Audio sourceQ3 outputs speech, sound effects and music natively, on by default in the API, with a speech-only modethree tracks in one pass
- Per-character bindingNot documented by the vendor (as of 2026-09-22)no public description
- Audio sourceThe accepted pairs are video with audio, video with text, image with audio and image with textsupplied or read from text
- Languagessync-3 supports 95 or more languages, described as the same coverage as the models before itpublished as a count
- Length per call on the free tierA free account runs one generation a month with a fifteen-second ceiling, with paid limits following the plan
5Sources
Every cell above is read from the documentation each vendor publishes, on the dates carried by the two model notes. The whole register on one column is on the field note; entries grouped by the shape of their answer are on the routes. Other pairs: Synthesia and sync-3, SceneMixer and Kling AI.