Sentioscope

Speech and voice controls, as each vendor documents them

sync-3: movement matched to the audio it is given

Lip movement is matched to the audio the model is given. Where the waveform is an input rather than an output, naming the driver costs the vendor nothing and tells a production exactly which half it controls. As of 2026-09-22.

A product whose entire job is this one columnWhere the waveform is an input rather than an output, naming the driver costs the vendor nothing. The interesting questions move downstream: how well, on what material, and at what cost.Audio inA recording you ownSets the timingDriver stated plainlyText inA scriptSynthesised, then matchedSame role assumed, notstatedPicture inFootage or a stillGets its mouth movedWhere the work landsOne sync-3 generationTwo of the four accepted pairs rely on an unnamed voice
Fig. 1 A free generation a month with a fifteen-second ceiling is enough to try the alignment on a production's own footage.
sync-3 on lip-sync, statement by statement. Read from the vendor's model documentation on 2026-09-22.
What the documentation settlesWhat it leaves to a take
Lip movement is matched to the audio the model is givenHow the match behaves when the audio is noisy or overlapping
The accepted pairs are video with audio, video with text, image with audio and image with textWhether the text variants align as well as the audio ones
A free account runs one generation a month with a fifteen-second ceilingWhat the alignment looks like on a production's own footage

Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.

1A product whose whole job is this column

Most entries here mention mouths in passing. This one exists to move them, so the driver statement is not a concession but a description of the product. The interesting questions move downstream: how well, on what material, and at what cost.

The free tier is the practical answer to the first of those. One generation a month with a fifteen-second ceiling is enough to try the alignment on real footage, which is exactly the test a page cannot substitute for.

2The text variants are a quieter claim

Two of the four accepted pairs take text, which means speech is synthesised somewhere before the mouth is matched. The driver statement is about audio, so a reader has to assume the synthesised audio plays the same role, and the page does not confirm it.

Recorded as a driver named for the audio routes. Whether the text routes behave identically is testable and unstated, and it matters because those routes are the ones a team with no recording session will reach for.

3The others that name what the mouth is following

Four more entries name a driver. Two are audio-driven like this one, one drives the mouth from a script and attaches a framing condition, and one makes the sound itself.

  • Audio source
    The accepted pairs are video with audio, video with text, image with audio and image with textsupplied or read from textSync, sync-3 model documentation / recorded 2026-09-22
  • Length per call on the free tier
    A free account runs one generation a month with a fifteen-second ceiling, with paid limits following the planSync, sync-3 model documentation / recorded 2026-09-22
  • Languages
    sync-3 supports 95 or more languages, described as the same coverage as the models before itpublished as a countSync, sync-3 model documentation / recorded 2026-09-22

4Sources

Read from the model documentation at sync.so on 2026-09-22. The same column across every entry is on lip-sync; everything this vendor publishes about speech is on sync-3. What counts as documented is on how read.