Sentioscope

Speech and voice controls, as each vendor documents them

Hedra: the audio is the input, not the output

Hedra does not make the sound. Avatar videos are driven by audio a production supplies, and the guide says the audio normally determines the video length. Omit the duration and the result follows the source recording. As of 2026-09-22.

When the audio is the input, the schedule runs backwardsAvatar videos are driven by audio that a production supplies, and omitting the duration makes the result follow the source recording. Casting, the session and a delivered file all sit in front of the first frame.Cast and recordA performer, asession and a filethat has to existfirst.A deliveredtrackApprove the readThe performance issigned off before aframe is rendered.An approved takeDrive the pictureThe character in theimage lip-syncs andmoves to that audio.
Fig. 1 The compensation for a longer front end is an approval that happens before any generation is billed.
Hedra on audio source, statement by statement. Read from the vendor's Character 3 model page and avatar video guide on 2026-09-22.
What the documentation settlesWhat it leaves to a take
Avatar videos are driven by audio, and the character lip-syncs to the audio providedHow closely a face tracks an overlapping or difficult read
The audio normally determines the video lengthWhether a ten-minute maximum is usable as one continuous piece
Speech can be generated inline by naming a voice id instead of uploading a trackWhether an inline voice lands where a recorded one would have

Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.

1A required input reverses the order of the work

Where sound comes out of the model, a shot can be attempted, heard and revised inside one pass. Where sound goes in, none of that is available until a recording exists. Casting, a session and a delivered file all sit upstream of the first frame, which lengthens the front of the schedule and shortens the part that involves rendering.

The compensation is an approval nobody else here offers. A director signs off the performance before a single generation is billed, and what comes back afterwards is a picture problem rather than a performance one.

2Length stops being a parameter and becomes a property of the take

Saying that the audio normally determines the video length removes duration from the request and puts it in the recording. That is convenient until a scene has to fit a slot, because trimming the picture now means re-cutting the audio and generating again rather than changing a number.

It also makes the pacing of the read the pacing of the shot. A performer who pauses for effect is buying frames, and a production working to a fixed episode length has to budget those pauses at the recording session, where they are cheap to change.

3The other entries that are handed a recording

Two more entries here take the sound as an input and fit a picture to it. Between them they differ on what else may ride alongside the audio, and on what the vendor will say about language.

  • sync-3 — supplied, or read from text.
  • Wan2.2-S2V — supplied; the model is audio-driven.
  • Audio source
    Avatar videos are driven by audio, and the character in the image will lip-sync and move to the audio providedhanded to the model rather than made by itHedra, avatar video guide / recorded 2026-09-22
  • Shot length set by the take
    The audio normally determines the video length, and omitting the duration follows the source audioHedra, avatar video guide / recorded 2026-09-22
  • Voice source
    Speech can be generated inline instead of uploaded, by naming a voice id drawn from the voices endpointa catalogue behind an endpointHedra, avatar video guide / recorded 2026-09-22

4Sources

Read from the Character 3 model page and avatar video guide at hedra.com on 2026-09-22. The same column across every entry is on audio source; everything this vendor publishes about speech is on Hedra. What counts as documented is on how read.