Hedra: the mouth follows the audio that was provided
The guide says the character in the image will lip-sync and move to the audio provided. That is a driver rather than a capability claim: the waveform is the instruction, and the picture is what follows it. As of 2026-09-22.
| What the documentation settles | What it leaves to a take |
|---|---|
| The character in the image will lip-sync and move to the audio provided | How closely the face tracks overlapping or crowded speech |
| Avatar videos are driven by audio, which is a required input | Whether a synthesised inline voice aligns as well as a recording |
| The audio normally determines the video length | Whether alignment holds over a ten-minute maximum |
Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.
1Naming the driver is worth more than naming the feature
A production reading that movement follows supplied audio knows what to do: approve the recording, then treat everything afterwards as a picture problem. A production reading only that lip-sync exists knows nothing about which half to fix when it looks wrong.
The sentence also promises less than it appears to. Following a waveform is the mechanism, not the quality, and nothing here says how the face behaves on plosives, on laughter, or on two voices arriving at once.
2Movement, not only mouths
The wording covers lip-sync and movement together, which is broader than most of this column. It implies head and body respond to the audio as well, and that is where audio-driven products usually look either impressive or uncanny.
Nothing published separates the two behaviours or says whether one can be reduced. A production that wants a still face and a moving mouth has no documented lever, and the register records the combined claim rather than splitting it.
3The others that name what the mouth is following
Four more entries name a driver rather than a feature. Two of those are also audio-driven, one drives the mouth from a script, and one makes the sound itself.
- MiniMax — whoever is on screen.
- sync-3 — the audio it is given.
- Synthesia — the spoken content, framed close.
- Wan2.2-S2V — the audio input.
- Audio sourceAvatar videos are driven by audio, and the character in the image will lip-sync and move to the audio providedhanded to the model rather than made by it
- Shot length set by the takeThe audio normally determines the video length, and omitting the duration follows the source audio
- LanguagesDescribed as text, image and audio to video with full multi-language support, and a maximum duration of ten minutesa claim carrying no list
4Sources
Read from the Character 3 model page and avatar video guide at hedra.com on 2026-09-22. The same column across every entry is on lip-sync; everything this vendor publishes about speech is on Hedra. What counts as documented is on how read.