Sentioscope

Speech and voice controls, as each vendor documents them

sync-3: four accepted pairs, and a free-tier ceiling

The model documentation lists four accepted pairs: video with audio, video with text, image with audio and image with text. Sound is therefore supplied or synthesised from text on the call, never produced alongside an invented picture. As of 2026-09-22.

Four accepted pairs, and four production shapesVideo with audio, video with text, image with audio and image with text. The first is a repair on existing footage, the second a talking portrait, and the text variants move speech synthesis inside this call.Video and audioFootage plus a trackRepair the mouthsLatest stage in apipelineVideo and textFootage plus a scriptSynthesise, thenmatchNo recordingsession neededImage and audioA still plus a trackAnimate a portraitNo footage neededat allImage and textA still plus a scriptSynthesise andanimateEverything insideone callOne sync-3 generationWhat each pair assumes already exists
Fig. 1 Listing all four makes the product's position explicit: it operates on material rather than inventing any.
sync-3 on audio source, statement by statement. Read from the vendor's model documentation on 2026-09-22.
What the documentation settlesWhat it leaves to a take
The accepted pairs are video with audio, video with text, image with audio and image with textWhich pair gives the steadiest result for a dialogue scene
Lip movement is matched to the audio the model is givenHow the match behaves when the supplied audio is noisy or overlapping
A free account runs one generation a month with a fifteen-second ceilingWhat the paid ceilings are, since they are described as following the plan

Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.

1Four pairs describe four different production shapes

Video with audio is a repair job on footage that already exists. Image with audio is a talking portrait. The two text variants move the speech synthesis inside this call instead of upstream of it. Listing all four makes the product's position explicit: it operates on material rather than inventing it.

For a drama pipeline that placement matters more than any parameter. This is a stage, and a stage has an input that something else must produce, which is a scheduling fact rather than a limitation.

2A published free ceiling is a real number in a field of vague ones

One generation a month with a fifteen-second ceiling is a small allowance and an unusually concrete one. It is enough to establish whether the alignment is acceptable on a production's own footage, which is exactly the question a page cannot answer.

The paid limits are described as following the plan rather than published as figures, so the register stops where the documentation does. A team that needs the ceiling in writing before committing has to ask for it.

3The other entries that are handed a recording

Two more entries take the sound as an input. Both of those are avatar or portrait products; this one also accepts existing footage, which puts it later in the pipeline.

  • Hedra — supplied; audio is required.
  • Wan2.2-S2V — supplied; the model is audio-driven.
  • Audio source
    The accepted pairs are video with audio, video with text, image with audio and image with textsupplied or read from textSync, sync-3 model documentation / recorded 2026-09-22
  • Length per call on the free tier
    A free account runs one generation a month with a fifteen-second ceiling, with paid limits following the planSync, sync-3 model documentation / recorded 2026-09-22
  • Languages
    sync-3 supports 95 or more languages, described as the same coverage as the models before itpublished as a countSync, sync-3 model documentation / recorded 2026-09-22

4Sources

Read from the model documentation at sync.so on 2026-09-22. The same column across every entry is on audio source; everything this vendor publishes about speech is on sync-3. What counts as documented is on how read.