Sentioscope

Speech and voice controls, as each vendor documents them

Vidu and sync-3: three tracks made, or one matched

Vidu documents three audio tracks generated natively and on by default, with a speech-only mode beside them. sync-3 matches lip movement to audio it is given. One invents a soundscape; the other repairs a performance. As of 2026-09-22.

A soundscape invented, against a performance repairedOne returns a finished mix unless the call says otherwise, so silence has to be asked for. The other returns nothing until a track is handed over, so silence is the default and sound is the input.ViduSpeech, effects, musicAll three by defaultSpeech-only modedocumentedsync-3Your recordingMouths matched to itDriver named plainlyWhat differsWho makes the soundAnd who fixes the faceGenerator against a stageSpeech on screen, two waysOne is nearly a deliverable, one is applied to one
Fig. 1 The usual objection to generated audio is the music rather than the speech, and only one of these documents a way to separate them.
Controls documented, model by modelFilled where the model documents that control, hollow where nothing is published about it.Controls documented, model by modelVidusync-3Where they partAudio sourceAudio source — Vidu: With the picture, speech available on its ownAudio source — sync-3: Supplied, or read from textAudio source — Where they part: Different answersVoice sourceVoice source — Vidu: Nothing documented as an inputVoice source — sync-3: Nothing documented as an inputVoice source — Where they part: Same answerPer-character bindingPer-character binding — Vidu: Not documented by the vendorPer-character binding — sync-3: Not documented by the vendorPer-character binding — Where they part: Same answerLanguagesLanguages — Vidu: Not documented by the vendorLanguages — sync-3: A count, ninety-five or moreLanguages — Where they part: Only sync-3 answersLip-syncLip-sync — Vidu: Not documented by the vendorLip-sync — sync-3: The audio it is givenLip-sync — Where they part: Only sync-3 answers
Fig. 2 Filled where the model documents that control, hollow where nothing is published about it.
Models documenting each controlHow many models document each control. A hollow column is a statement about documentation, not capability.Models documenting each controlVidu2 of 5sync-34 of 5Where they part5 of 5
Fig. 3 How many models document each control. A hollow column is a statement about documentation, not capability.
Vidu and sync-3 on the five fields this register keeps, with the gap on each. Recorded 2026-09-22.
FieldVidusync-3Where they part
Audio sourceWith the picture, speech available on its ownSupplied, or read from textDifferent answers
Voice sourceNothing documented as an inputNothing documented as an inputSame answer
Per-character bindingNot documented by the vendorNot documented by the vendorSame answer
LanguagesNot documented by the vendorA count, ninety-five or moreOnly sync-3 answers
Lip-syncNot documented by the vendorThe audio it is givenOnly sync-3 answers

Inclusion rule. Two entries are given a page together when at least one column puts them at opposite grades of answer. Pairs that agree on every column, or that are both blank throughout, do not get a page. Order. Fixed field order, identical on every side-by-side page.

1A default soundscape against a supplied one

One of these returns a finished mix unless the call says otherwise, which makes silence something a production has to ask for. The other returns nothing until a track is handed over, so silence is the default and sound is the input.

Neither is better and they belong at different points in a pipeline. The first is a generator whose output is nearly a deliverable; the second is a stage applied to something already made.

2The speech-only mode is the interesting overlap

Keeping the dialogue and dropping the bed is documented on only one entry in this register, and it is the generator. That matters because the usual objection to generated audio is the music, not the speech, and a mode that separates them removes the objection without removing the feature.

The other entry never has that problem, because the bed was never generated. What it has instead is a dependency: something upstream must produce the dialogue, and that something is not described on its page.

3One names a driver, and the other names nothing about mouths

Matching lip movement to supplied audio is stated plainly on one page. The other describes three tracks and never says what happens to a face, which is the pattern for native-audio entries throughout this register.

Set side by side that is the clearest argument for keeping the lip-sync column at all. Two products that both put speech on screen give a reader completely different amounts to plan with.

4Each of them on its own

The column this pair was chosen for is audio source, and each entry has a page of its own on it. The full row for either, all five columns with the wording behind each cell, is on its model note.

  • Audio source
    Q3 outputs speech, sound effects and music natively, on by default in the API, with a speech-only modethree tracks in one passVidu, API reference / recorded 2026-09-12
  • Per-character binding
    Not documented by the vendor (as of 2026-09-22)no public descriptionVidu, API reference / recorded 2026-09-12
  • Audio source
    The accepted pairs are video with audio, video with text, image with audio and image with textsupplied or read from textSync, sync-3 model documentation / recorded 2026-09-22
  • Languages
    sync-3 supports 95 or more languages, described as the same coverage as the models before itpublished as a countSync, sync-3 model documentation / recorded 2026-09-22
  • Length per call on the free tier
    A free account runs one generation a month with a fifteen-second ceiling, with paid limits following the planSync, sync-3 model documentation / recorded 2026-09-22

5Sources

Every cell above is read from the documentation each vendor publishes, on the dates carried by the two model notes. The whole register on one column is on the field note; entries grouped by the shape of their answer are on the routes. Other pairs: Synthesia and sync-3, SceneMixer and Kling AI.