Sentioscope

Speech and voice controls, as each vendor documents them

Hedra and Synthesia: a waveform, or the script itself

Both name what the mouth follows, which only five entries here manage. One follows the audio a production supplies; the other generates lip sync and expression from the script, and adds a condition about framing. As of 2026-09-22.

Following a waveform, against following wordsWhere the driver is a recording, a performer's pauses and emphasis reach the screen. Where the driver is the script, the model decides those, and the same words always produce roughly the same reading.HedraSynthesiaWhat the mouth followsThe audio providedThe spoken content of the scriptWho sets the timingThe performer, in the sessionThe model, from the wordsA weak case namedNoneWide framing performs less wellFitting a slotRe-recordRewriteTwo products that look alike and have different customers
Fig. 1 For drama the first preserves a performance and the second removes one; for a presentation the second is cheaper and adequate.
Controls documented, model by modelFilled where the model documents that control, hollow where nothing is published about it.Controls documented, model by modelHedraSynthesiaWhere they partAudio sourceAudio source — Hedra: Supplied; audio is requiredAudio source — Synthesia: From the script, or uploadedAudio source — Where they part: Different answersVoice sourceVoice source — Hedra: A track, or a voice id from the endpointVoice source — Synthesia: A catalogue voice, or a cloned oneVoice source — Where they part: Different answersPer-character bindingPer-character binding — Hedra: Not documented by the vendorPer-character binding — Synthesia: Not documented by the vendorPer-character binding — Where they part: Same answerLanguagesLanguages — Hedra: A claim with no namesLanguages — Synthesia: A catalogue with codesLanguages — Where they part: Different answersLip-syncLip-sync — Hedra: The supplied audioLip-sync — Synthesia: The spoken content, framed closeLip-sync — Where they part: Different answers
Fig. 2 Filled where the model documents that control, hollow where nothing is published about it.
Models documenting each controlHow many models document each control. A hollow column is a statement about documentation, not capability.Models documenting each controlHedra4 of 5Synthesia4 of 5Where they part5 of 5
Fig. 3 How many models document each control. A hollow column is a statement about documentation, not capability.
Hedra and Synthesia on the five fields this register keeps, with the gap on each. Recorded 2026-09-22.
FieldHedraSynthesiaWhere they part
Audio sourceSupplied; audio is requiredFrom the script, or uploadedDifferent answers
Voice sourceA track, or a voice id from the endpointA catalogue voice, or a cloned oneDifferent answers
Per-character bindingNot documented by the vendorNot documented by the vendorSame answer
LanguagesA claim with no namesA catalogue with codesDifferent answers
Lip-syncThe supplied audioThe spoken content, framed closeDifferent answers

Inclusion rule. Two entries are given a page together when at least one column puts them at opposite grades of answer. Pairs that agree on every column, or that are both blank throughout, do not get a page. Order. Fixed field order, identical on every side-by-side page.

1Following a waveform against following words

Where the driver is a recording, a performer's timing reaches the screen: pauses, emphasis and breath are all in the file. Where the driver is the script, the model decides those things, and the same words will always produce roughly the same reading.

For drama the first preserves a performance and the second removes one. For a presentation the second is cheaper and entirely adequate, which is why these two products have different customers despite looking alike.

2Only one of them publishes a weak case

Saying performance improves with the avatar framed closer names the shot where alignment struggles, which is a constraint on coverage rather than a tip. The audio-driven entry names its driver and says nothing about when it holds less well.

Both sentences are worth having and the second kind is rarer. A vendor willing to publish where something fails has given a production more than one willing only to publish that it works.

3Length comes from opposite places

One says the audio normally determines the video length, so duration lives in the recording and a slot is fitted by re-recording. The other drives everything from the script, so a rewrite is the lever and a length change is a text change.

Neither publishes a language list. One claims full multi-language support; the other publishes a code on every voice, which is the widest gap between these two rows and the one that decides which markets are reachable.

4Each of them on its own

The column this pair was chosen for is lip-sync, and each entry has a page of its own on it. The full row for either, all five columns with the wording behind each cell, is on its model note.

  • Audio source
    Avatar videos are driven by audio, and the character in the image will lip-sync and move to the audio providedhanded to the model rather than made by itHedra, avatar video guide / recorded 2026-09-22
  • Shot length set by the take
    The audio normally determines the video length, and omitting the duration follows the source audioHedra, avatar video guide / recorded 2026-09-22
  • Languages
    Described as text, image and audio to video with full multi-language support, and a maximum duration of ten minutesa claim carrying no listHedra, Character 3 model page / recorded 2026-09-22
  • Lip-sync
    Lip sync and facial expressions are generated from the spoken contentdriven by the scriptSynthesia, create an avatar / recorded 2026-09-22
  • A framing condition attached to lip-sync
    Lip sync performs best when the avatar is framed closer in the scene rather than positioned far from the cameraSynthesia, create an avatar / recorded 2026-09-22

5Sources

Every cell above is read from the documentation each vendor publishes, on the dates carried by the two model notes. The whole register on one column is on the field note; entries grouped by the shape of their answer are on the routes. Other pairs: Wan 3.0 and Wan2.2-S2V, Veo and Hedra.