Sentioscope

Speech and voice controls, as each vendor documents them

Veo and Hedra: sound out of one, into the other

Veo documents audio generated with the video and stops. Hedra requires audio as an input and says the character lip-syncs and moves to it. Two sparse entries, sparse in opposite places. As of 2026-09-22.

Two thin pages, thin in opposite placesA sentence saying audio arrives with the video removes a dubbing stage and says nothing about casting. A sentence saying audio is required puts a session on a schedule and says nothing about what comes back.VeoHedraWhat the one sentence settlesNo dubbing stage is neededA recording session is neededWhat it leaves openWhose voice, and in what languageHow closely the face followsThe mouthNot mentionedFollows the audio providedOwning a voiceNot possible from the pagePossible, because you supply itFor a serial, being able to own the voice decides it
Fig. 1 A required input is a task with a date; a produced output is a hope with a date.
Controls documented, model by modelFilled where the model documents that control, hollow where nothing is published about it.Controls documented, model by modelVeoHedraWhere they partAudio sourceAudio source — Veo: With the pictureAudio source — Hedra: Supplied; audio is requiredAudio source — Where they part: Different answersVoice sourceVoice source — Veo: Not documented by the vendorVoice source — Hedra: A track, or a voice id from the endpointVoice source — Where they part: Only Hedra answersPer-character bindingPer-character binding — Veo: Not documented by the vendorPer-character binding — Hedra: Not documented by the vendorPer-character binding — Where they part: Same answerLanguagesLanguages — Veo: Not documented by the vendorLanguages — Hedra: A claim with no namesLanguages — Where they part: Only Hedra answersLip-syncLip-sync — Veo: Not documented by the vendorLip-sync — Hedra: The supplied audioLip-sync — Where they part: Only Hedra answers
Fig. 2 Filled where the model documents that control, hollow where nothing is published about it.
Models documenting each controlHow many models document each control. A hollow column is a statement about documentation, not capability.Models documenting each controlVeo1 of 5Hedra4 of 5Where they part5 of 5
Fig. 3 How many models document each control. A hollow column is a statement about documentation, not capability.
Veo and Hedra on the five fields this register keeps, with the gap on each. Recorded 2026-09-22.
FieldVeoHedraWhere they part
Audio sourceWith the pictureSupplied; audio is requiredDifferent answers
Voice sourceNot documented by the vendorA track, or a voice id from the endpointOnly Hedra answers
Per-character bindingNot documented by the vendorNot documented by the vendorSame answer
LanguagesNot documented by the vendorA claim with no namesOnly Hedra answers
Lip-syncNot documented by the vendorThe supplied audioOnly Hedra answers

Inclusion rule. Two entries are given a page together when at least one column puts them at opposite grades of answer. Pairs that agree on every column, or that are both blank throughout, do not get a page. Order. Fixed field order, identical on every side-by-side page.

1Two thin pages, and the thinness lands differently

Neither of these publishes much. The difference is which question each one answers. A sentence saying audio arrives with the video tells a production a dubbing stage is unnecessary and nothing about casting. A sentence saying audio is required tells a production to book a session and nothing about what comes back.

For planning, the second is the more actionable. A required input is a task on a schedule; a produced output is a hope with a date on it.

2Only one of them addresses the mouth

The audio-driven entry names its driver because the driver is the input, which costs the vendor nothing to say. The native-audio entry never mentions mouths, which is the common position among entries that make sound and picture together.

That pattern holds across this register and is worth stating plainly: a production that needs a published statement about lip movement is mostly choosing between products that are handed a recording.

3Both leave the cast entirely open

Neither publishes a language list, and neither describes a voice attaching to a character. One of them at least lets a production supply the voice, which makes consistency an archival problem rather than an impossible one.

For a serial that is the deciding difference. A voice a production owns can be repeated in episode nine; a voice the model invents on each call cannot be, and no sentence on either page says otherwise.

4Each of them on its own

The column this pair was chosen for is audio source, and each entry has a page of its own on it. The full row for either, all five columns with the wording behind each cell, is on its model note.

  • Audio source
    Audio is generated with the video rather than added afterwardsgenerated with the pictureGoogle, Veo documentation / recorded 2026-09-12
  • Audio source
    Avatar videos are driven by audio, and the character in the image will lip-sync and move to the audio providedhanded to the model rather than made by itHedra, avatar video guide / recorded 2026-09-22
  • Shot length set by the take
    The audio normally determines the video length, and omitting the duration follows the source audioHedra, avatar video guide / recorded 2026-09-22
  • Languages
    Described as text, image and audio to video with full multi-language support, and a maximum duration of ten minutesa claim carrying no listHedra, Character 3 model page / recorded 2026-09-22

5Sources

Every cell above is read from the documentation each vendor publishes, on the dates carried by the two model notes. The whole register on one column is on the field note; entries grouped by the shape of their answer are on the routes. Other pairs: Vidu and sync-3, Synthesia and sync-3.