What Hedra documents about speech
Hedra (hedra.com) inverts the usual arrangement: audio is a required input, the character in the image lip-syncs and moves to whatever is supplied, and the length of the take sets the length of the shot, up to ten minutes. As of 2026-09-22.
| Field | What the vendor documents |
|---|---|
| Audio source | Supplied; audio is a required input |
| Languages | Full multi-language support claimed, none named |
| Voice source | An uploaded track, or a voice id from the voices endpoint |
| Per-character binding | Not documented |
| Lip-sync | The character in the image lip-syncs and moves to the audio provided |
Inclusion rule. Fields are the same five for every model here; a field the vendor does not document is recorded as not documented rather than inferred from generated output. Order. Fixed field order, identical on every model page.
1When the audio is the input, alignment stops being a promise
The guide states that avatar videos are driven by audio and that the character in the image will lip-sync and move to the audio provided. Synchronisation is not a feature the model might achieve; it is the task it was handed. That makes this the plainest lip-sync statement in the register, and it is plain because of the architecture rather than the copywriting.
The trade is that a performance has to exist first. A production on this route is casting and recording before it generates anything, which is an older order of work than prompting a scene and hearing what comes back.
2Length decided by the take, not by a parameter
The audio normally determines the video length, and omitting the duration follows the source audio. For a register that spends most of its time on ceilings measured in single-figure seconds, a ten-minute maximum is a different world, and a line that runs long produces a longer shot instead of a rushed read.
It moves the timing problem back into writing and recording, where a director can hear it. Nothing has to be trimmed to fit a four-second window, because the window follows the words.
3Voices are available, and the language claim carries no list
Speech can be generated inline by naming a voice id from the voices endpoint, so a production need not arrive with a recording at all. What the model page then claims is full multi-language support without naming a single language, which is the count-free version of a gap several entries here share.
Nothing is published about holding a voice to a character across shots. With audio supplied per call, consistency is whatever a production's own choices make it, which is workable and unstated.
4This entry, one column at a time
Each of these stays inside a single field: what Hedra puts there, what the wording settles, and what it leaves for a take to answer.
- Audio source — supplied; audio is required.
- Voice source — a track, or a voice id from the endpoint.
- Per-character binding — not documented by the vendor.
- Languages — a claim with no names.
- Lip-sync — the supplied audio.
5Read against another entry
Each of these puts Hedra beside one other entry on a column where the two land at opposite grades of answer.
- Veo and Hedra — on audio source.
- D-ID and Hedra — on audio source.
- Hedra and Synthesia — on lip-sync.
6Where this model sits against the rest
| Model | Audio source |
|---|---|
| D-ID | A text script read aloud, or an audio url supplied |
| Hedra | Supplied; audio is a required input |
| HeyGen | A text to speech endpoint, apart from the picture |
| Kling AI | Generated with the picture, VIDEO 3.0 |
| LTX Studio | Generated jointly with the picture, plus audio-to-video |
| Luma Ray | Not documented; the video reference does not mention sound |
| MiniMax | Native speech, generated with the picture |
| PixVerse | Supplied to a speech endpoint, or read from text there |
| Runway | Not documented where voices are defined |
| SceneMixer | Generated with the picture, in the language set for the project |
| Sora 2 | Generated with the picture; video and audio are both listed as output |
| sync-3 | Supplied, or read from text on the call |
| Synthesia | Synthesised from the script, or an uploaded recording |
| Veo | Generated with the video rather than added afterwards |
| Vidu | Q3 generates speech, effects and music natively, on by default |
| Wan 3.0 | On by default; false returns a file with no audio track |
| Wan2.2-S2V | Supplied; the card describes the model as audio-driven |
Inclusion rule. Models whose vendor documents something about generated speech. A model with nothing published on the point is still listed, with the cell marked hollow. Order. Alphabetical by model name.
| Model | Languages |
|---|---|
| D-ID | A language field beside the voice; no list on the reference |
| Hedra | Full multi-language support claimed, none named |
| HeyGen | Thirty or more, counted on the translation endpoint |
| Kling AI | Five, not named |
| LTX Studio | Not documented |
| Luma Ray | Not documented |
| MiniMax | Not documented by the vendor |
| PixVerse | Multiple, with speech, singing and advertisements named as types |
| Runway | Not documented |
| SceneMixer | 15, named, with Cantonese a 16th for dialogue only |
| Sora 2 | Not documented |
| sync-3 | Ninety-five or more, counted rather than named |
| Synthesia | A catalogue, each voice carrying a language name and code |
| Veo | Not documented |
| Vidu | Not documented |
| Wan 3.0 | Not documented |
| Wan2.2-S2V | Not documented |
Inclusion rule. Models whose vendor documents something about generated speech. A model with nothing published on the point is still listed, with the cell marked hollow. Order. Alphabetical by model name.
| Model | Voice source |
|---|---|
| D-ID | A voice id from one of five named speech providers |
| Hedra | An uploaded track, or a voice id from the voices endpoint |
| HeyGen | A clone, instant from one recording or professional from 20 minutes or more |
| Kling AI | Not documented by the vendor |
| LTX Studio | Not documented |
| Luma Ray | Not documented |
| MiniMax | Reference audio, 15 s total across 3 clips |
| PixVerse | Built-in voices, or a custom voice from a supplied sample |
| Runway | An audio sample of 10 seconds to 5 minutes, or a description in words |
| SceneMixer | Preset voice library, or the user's own recordings |
| Sora 2 | Not documented |
| sync-3 | Not documented on the model page |
| Synthesia | The catalogue, or a cloned voice with a language list of its own |
| Veo | Not documented |
| Vidu | Not documented |
| Wan 3.0 | Reference audio, 15 seconds in total, WAV or MP3 up to 15 MB |
| Wan2.2-S2V | Whatever track is handed over; no catalogue |
Inclusion rule. Models whose vendor documents something about generated speech. A model with nothing published on the point is still listed, with the cell marked hollow. Order. Alphabetical by model name.
7Sources
Taken from the Character 3 model page and avatar video guide at hedra.com on 2026-09-22. Every entry against the same five fields sits on the speech controls page; the rule behind each field is on how read. The entries either side of this one: HeyGen, Kling AI.
- Audio sourceAvatar videos are driven by audio, and the character in the image will lip-sync and move to the audio providedhanded to the model rather than made by it
- Shot length set by the takeThe audio normally determines the video length, and omitting the duration follows the source audio
- Voice sourceSpeech can be generated inline instead of uploaded, by naming a voice id drawn from the voices endpointa catalogue behind an endpoint
- LanguagesDescribed as text, image and audio to video with full multi-language support, and a maximum duration of ten minutesa claim carrying no list