Hedra: the audio is the input, not the output
Hedra does not make the sound. Avatar videos are driven by audio a production supplies, and the guide says the audio normally determines the video length. Omit the duration and the result follows the source recording. As of 2026-09-22.
| What the documentation settles | What it leaves to a take |
|---|---|
| Avatar videos are driven by audio, and the character lip-syncs to the audio provided | How closely a face tracks an overlapping or difficult read |
| The audio normally determines the video length | Whether a ten-minute maximum is usable as one continuous piece |
| Speech can be generated inline by naming a voice id instead of uploading a track | Whether an inline voice lands where a recorded one would have |
Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.
1A required input reverses the order of the work
Where sound comes out of the model, a shot can be attempted, heard and revised inside one pass. Where sound goes in, none of that is available until a recording exists. Casting, a session and a delivered file all sit upstream of the first frame, which lengthens the front of the schedule and shortens the part that involves rendering.
The compensation is an approval nobody else here offers. A director signs off the performance before a single generation is billed, and what comes back afterwards is a picture problem rather than a performance one.
2Length stops being a parameter and becomes a property of the take
Saying that the audio normally determines the video length removes duration from the request and puts it in the recording. That is convenient until a scene has to fit a slot, because trimming the picture now means re-cutting the audio and generating again rather than changing a number.
It also makes the pacing of the read the pacing of the shot. A performer who pauses for effect is buying frames, and a production working to a fixed episode length has to budget those pauses at the recording session, where they are cheap to change.
3The other entries that are handed a recording
Two more entries here take the sound as an input and fit a picture to it. Between them they differ on what else may ride alongside the audio, and on what the vendor will say about language.
- sync-3 — supplied, or read from text.
- Wan2.2-S2V — supplied; the model is audio-driven.
- Audio sourceAvatar videos are driven by audio, and the character in the image will lip-sync and move to the audio providedhanded to the model rather than made by it
- Shot length set by the takeThe audio normally determines the video length, and omitting the duration follows the source audio
- Voice sourceSpeech can be generated inline instead of uploaded, by naming a voice id drawn from the voices endpointa catalogue behind an endpoint
4Sources
Read from the Character 3 model page and avatar video guide at hedra.com on 2026-09-22. The same column across every entry is on audio source; everything this vendor publishes about speech is on Hedra. What counts as documented is on how read.